← Knowledge

Context language models

How model-editable live context works, what the paper's benchmarks report, and the limits of editable context and approximate cache reuse.

Context language models (CLMs) are language-model agents that can edit the context supplied to their next model call. In the September 2026 paper, the authors implement this by exposing the live context as a writable file. An agent can remove old observations, rewrite a working summary, or maintain a compact task ledger using ordinary code. The harness synchronizes those edits with the next model input.

The proposal makes context management a behavior the model can learn. The authors report improvements on research and coding tasks using existing models, then study instruction optimization, reinforcement learning, and a serving optimization for edited contexts. The results come from specific benchmark settings; they do not establish that unrestricted context editing is reliable or safe in every deployment.

Editing the next input

A conventional agent conversation grows by appending messages until its harness summarizes or trims the history. Some systems expose predefined tools for compaction or retrieval. CLMs instead let the model define the transformation itself through edits to a context file.

For example, an agent investigating a software failure might accumulate dozens of unsuccessful searches. It could replace those results with the queries already tried, preserve the relevant error verbatim, and update a list of remaining hypotheses. This is an illustrative use of the mechanism, not a reported experiment.

The important implementation detail is synchronization. An ordinary notes file affects the model only when something retrieves or injects it. The CLM context file determines the next live input. When the agent makes no edit, newly generated output is appended normally.

The paper also extends the design to multiple agents with separate context files. Reported behaviors include maintaining worker-status tables, inserting internal notes, and writing reusable functions to compact old search results.

What ContextBench measures

The authors introduce ContextBench to separate context-management problems from demanding domain knowledge. It contains four synthetic tasks:

  • Needle Retention: preserve selected information exactly while other material accumulates.
  • Sudoku Sketchpad: maintain an evolving board through small updates.
  • KV Store: offload and retrieve large values by key.
  • Log Triage: retain access to useful evidence within large logs.

Experiments use a 32,000-token context limit and vary input volume up to 24 times that limit. The tasks expose different weaknesses: summaries can lose exact details, append-only histories can retain obsolete state, and external storage alone does not remove material already occupying the live context.

These tests diagnose particular memory operations. Their simplicity helps isolate those operations, but limits what success says about open-ended work.

Results and their scope

On BrowseComp-Plus, a deep-research benchmark, the paper reports 59.4% accuracy for zero-shot CLM using Qwen3.6-27B with a 32K context budget. That is an 11.4% relative improvement, not an 11.4-percentage-point increase, over its strongest baseline. Estimated prefix-reuse compute was 21.5% lower.

On TerminalBench 2.1, CLM matched the strongest baseline while using about 29.5% fewer prefix-reuse FLOPs. On a ten-task EdgeBench subset, the Qwen model scored 44.6 versus 42.3 for summarization, using 179 versus 437 prefix-reuse PFLOPs per trial. Scores were best-of-three over runs lasting up to twelve hours. Adding subagents offered little additional benefit in that single-repository setting.

The multi-repository experiment used six agents and evaluated their changes on unseen downstream packages. Its reported 65% greater downstream speedup concerns that software-optimization setup at matched spend. It is not a general claim that agents become 65% faster.

FLOPs estimate arithmetic work. They do not directly measure API price, latency, or total operational cost. The authors explicitly account for recomputation after edits rather than assuming shorter context automatically means cheaper inference.

Learning a context-management policy

The paper tests two ways to improve how CLMs edit context.

Instruction optimization changes the guidance supplied to a model while keeping its weights fixed. A proposer examines training trajectories and suggests context-management skills. Candidates are selected on development data, with the final selection evaluated on a held-out test set.

Reinforcement learning changes the model weights. Each model-call segment receives an advantage derived from the complete trajectory's outcome. An additional efficiency term rewards cheaper trajectories only among successful attempts. The intent is to discourage deleting useful information merely to reduce computation.

For Qwen3.5-9B on BrowseComp-Plus, CLM accuracy rose from 28.8% before training to 42.5% afterward. The trained summarization baseline reached 42.1%. CLM used 1.34 PFLOPs per question versus 2.19 for that trained baseline. Before training, the smaller CLM performed worse than summarization, an important limit on the claim that more context control is automatically better.

Suffix cache reuse

Inference servers cache intermediate states to avoid recomputing an unchanged prompt prefix. Editing the middle of a context normally invalidates reuse from that point onward, even when later text remains unchanged.

The paper's Suffix Cache Reuse preserves cached states for surviving suffix tokens and recomputes inserted material. Those retained states are stale: they were calculated with the previous prefix. This is approximate reuse, not an exact recalculation of the edited context.

The authors report matched task performance with about 35% less server-side compute than standard SGLang in the tested BrowseComp-Plus configuration. This saving is reported separately from the main context-management comparisons. The experiment does not establish that stale states are harmless for every task.

Limits and safety

Editable context gives an agent another place to persist misleading instructions, including its own mistaken conclusions or material originating in a prompt injection. The paper explicitly identifies this risk and leaves defenses as future work.

A useful engineering distinction is between editable working context and an authoritative record of what actually happened. As a design inference, retaining an external event history and checking context edits against protected instructions would make changes more inspectable. The paper's benchmark gains do not verify such a production control system. This distinction also informs reliable agent systems.

CLMs also address a different problem from reading a large external document on demand. Retrieval controls what enters the context; live-context editing controls what remains and how it is represented. Both can be useful in the same agent. Agent memory considers these as different layers of state.

Sources

  • Paper, submitted September 29, 2026.
  • Full text, including benchmark settings, learning procedures, serving design, and limitations.
  • Implementation, linked by the authors.

Sources

  1. Context Language Models
  2. Context Language Models implementation

Connections

Related

Linked here

Suggest a correction ↗

Appearance