Weekly Paper Notes — one of the top picks from the 2026-10-03 CS paper digest. Area: NLP.

Authors: Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh (University of Washington, Meta Superintelligence Labs, MIT, Trillium Labs) arXiv: 2609.37725 · PDF · Code

TL;DR

A standard language model’s context only grows: each turn appends the model’s output, c_{t+1} = c_t ⊕ f(c_t). Anything smarter, such as compaction, offloading or summarisation, is done by the harness according to rules a human wrote. A Context Language Model (CLM) replaces that with c_{t+1} = f_CLM(c_t), where the model itself produces the next context.

The implementation is deliberately simple. The live context is mirrored to a file whose path appears in the system prompt. The model can leave it alone (output gets appended as usual) or edit it with ordinary Bash and Python, and whatever is in the file becomes the context for the next call. Several context files can coexist, so subagents and agent swarms come for free.

Applied zero-shot to existing models with a 32K budget, the approach beats every harness-defined and tool-based baseline the authors tried. On BrowseComp-Plus it scores 59.4%, an 11.4% relative gain over Codex-style summarisation with 21.5% fewer FLOPs. It matches summarisation on TerminalBench 2.1 with about 30% fewer FLOPs, scores 5% higher with 59% fewer FLOPs on 12-hour EdgeBench runs, and gets 65% more downstream speedup on a 24-hour, six-repository agent-swarm task. The paper also shows the strategy can be steered by one sentence of instruction, improved by evolving a skill document, and trained into weights with RL, and it adds a serving trick, Suffix Cache Reuse, that cuts server compute another 35%.

What problem is the paper actually attacking?

Long-horizon agents run out of context. A deep-research trajectory can run past a hundred tool calls and a repository-optimisation run can last twelve hours, so something has to decide what stays in the window.

The answers so far form a ladder of increasing model autonomy. At the bottom are harness-defined policies: Codex-style summarisation when the window fills, MEM1-style consolidation every turn, Context Folding. One step up, the model gets a fixed set of context actions: AutoCompact and Self-Compact let it choose when to compact, Context-as-a-Tool exposes compaction over a predefined region, ACM adds model-triggered offload and retrieval, and Sculptor lets it pick fragments to operate on. Recursive Language Models (RLMs) take another route and treat a long input as a REPL variable the model reads on demand, which handles what to read in but not how to manage what’s already in the live context.

Every rung keeps the action space human-defined. The authors set out to show that this is the constraint, and that a capable model with unrestricted access to its context will find better strategies than the menus it’s been given. They frame it explicitly as an instance of the Bitter Lesson.

The pilot study makes the case concrete. ContextBench has four synthetic tasks that isolate context management from reasoning: Needle Retention (keep particular strings verbatim while unrelated input streams past), Sudoku Sketchpad (keep a board current as moves arrive, which needs small in-place edits), and KV Store and Log Triage (offload large values and retrieve them exactly). At a 32K window and up to 24x context pressure, every fixed strategy fails somewhere. Summaries lose or invent needles and board cells. Methods without in-place editing have to rewrite the whole board for each move. Coding tools can offload a value to disk but can’t evict it from the live window. No baseline is perfect even on tasks this simple, and that’s the gap CLMs are meant to close.

The mechanism: context as a file

Figure 1 of the paper: (a) examples of emergent context edits, (b) results on BrowseComp-Plus and Software World, (c) steering by instruction and skill evolution, (d) RL with a success-gated efficiency advantage. Figure: The CLM recipe and its four results threads. Source: Shao et al., arXiv:2609.37725, Figure 1.

The formal change is the move from Eq. 1 to Eq. 2:

standard LM:  c_{t+1} = c_t ⊕ f_θ^LM(c_t)
CLM:          c_{t+1} = f_θ^CLM(c_t)

Eq. 2 contains every earlier approach as a special case: any compaction tool is a particular f. What’s new is that the model defines its context-management functions instead of being handed them.

The context-as-a-file implementation keeps this general without inventing a new interface. The model already knows how to edit files from code-agent training. Edits to the context file are synced to the live context and sent to the server for the next turn, and if the model makes no edit, new tokens are appended as normal, which keeps prefix caching intact on most turns. A multi-agent setup is just several context files, and spawning or killing a subagent is creating or deleting one.

The qualitative examples are the most interesting part of the paper. Running zero-shot, models invent context machinery nobody asked for. In one Erdős minimum-overlap run, the orchestrator keeps an in-context scoreboard of 21 subagents and updates it in place 163 times while holding its own context to 6–8K tokens. Another run invents a new chat role, notes, next to the template’s user and assistant. In a circle-packing run the model defines a compact_turns helper and calls it 37 times to fold old observations into pointers to a progress note. Others write loops that strip irrelevant search results or keep a ledger of untried ideas. A human-designed action menu wouldn’t have offered most of these.

The practical claim: does editing the middle of context ruin the cache?

The obvious objection is serving cost. Prefix caching only helps while the prompt is unchanged from the front, and an edit in the middle forces everything after it to be re-prefilled. A model that edits its context constantly could easily cost more than it saves.

The paper answers in two ways. First, it measures cost honestly with prefix-reuse FLOPs: prefill for every token from the first mismatch onward plus decode for new tokens, summed over the trajectory. Every efficiency number in the paper uses this metric under standard serving, so the gains above already include the re-prefill penalty. The CLM still comes out cheaper because it keeps the window small.

Suffix Cache Reuse compared with standard serving: after segment B is replaced by B′, standard serving re-prefills B′ and everything after it, while SCR re-prefills only B′ and reuses cached states for surviving tokens C. Figure: Standard prefix reuse versus Suffix Cache Reuse. Source: Shao et al., arXiv:2609.37725, Figure 4.

Second, it introduces Suffix Cache Reuse (SCR). When segment B becomes B′, SCR re-prefills only B′ and keeps the cached KV states for the surviving suffix C, even though those states were computed against the old prefix. That’s approximate, and the authors note the stale states sometimes help because they carry information from content that’s been deleted. On BrowseComp-Plus with Qwen3.6-27B, SCR matches standard SGLang accuracy at 65% of the compute. It also applies outside CLMs: chat templates that strip previous turns’ reasoning tokens cause the same mid-context mismatch, and SCR recovers that cost too. The appendix covers hybrid models that mix full and linear attention, where the recurrent state complicates reuse.

Results

Zero-shot coding and research. With Qwen3.6-27B, a 32K window and a 100-turn cap, the CLM sits on the Pareto frontier on all three benchmarks. BrowseComp-Plus is 59.4%, 11.4% relative above summarisation, with 21.5% fewer FLOPs than summarisation and 28.9% fewer than MEM1. TerminalBench 2.1 ties summarisation at 70% of its FLOPs. TBLite is 73.7% against 67.0% at 91% of the FLOPs. RLM, ACM and Self-Compact all land below the frontier.

Open discovery. On four AlphaEvolve problems using Claude 4.6 Sonnet, a general CLM agent given the evolutionary algorithm as in-context guidance gets the best score on all four, ahead of the specialised OpenEvolve workflow and an agentic OpenEvolve variant. The margin is up to 16.8% on Heilbronn and 3.0% on circle packing. On EdgeBench-10, twelve-hour single-repository optimisation, the Qwen CLM scores 44.6 at 179 PFLOPs per trial against summarisation’s 42.3 at 437. With Sonnet it scores 51.0 against 42.3. Up to five subagents add nothing there. On Software World, six GPT-5.6-Sol agents with 272K windows jointly optimise interdependent Python packages (requests, urllib3 and others) for over 24 hours and are scored on four unseen downstream consumers. The CLM swarm gets 65% more downstream speedup than a summarisation swarm with the same spend.

Steering and evolution. One sentence in the prompt (“compact to 4k tokens once you reach Y”, “compact at sub-question boundaries”, “back up on disk before you compact”) changes the context-management policy measurably, with no change to harness or weights. An evolved skill document improves held-out ContextBench accuracy by up to 35.9 points while lowering compute, whether the proposer is a stronger external model or the agent itself.

RL. The training recipe is stepwise GRPO. Every model call in a trajectory gets that trajectory’s outcome advantage, plus a success-gated efficiency advantage:

A_i^eff = clip((c̄_g − c_i) / c̄_g, −1, 1)   if τ_i succeeded
        = 0                                otherwise
A_i     = A_i^out + w_eff · A_i^eff

c_i is the trajectory’s prefix-reuse FLOPs and c̄_g is the mean over successful trajectories in the group. The gate is the important design choice: rewarding edit counts or tokens removed would teach the model to throw away information, so efficiency only reorders trajectories that already succeeded. Training Qwen3.5-9B on OpenResearcher, the CLM goes from 28.8% to 42.5% on BrowseComp-Plus at 1.34 PFLOPs per question, while a summarisation harness trained with the same recipe reaches 42.1% at 2.19. The untrained 9B CLM started six points behind summarisation, so smaller models need training before full context control pays off.

Why this matters

In practice this moves context engineering from harness code into the model. Today’s agent frameworks carry a lot of hand-built compaction logic: thresholds, summary prompts, memory tiers, retrieval hooks. This paper argues that a capable model given a file and Bash will do at least as well and produce strategies you wouldn’t have written. A team maintaining that logic should consider replacing it with one context file and a short skill document, then tuning the document.

It also turns context management into a training target. Once the policy is model behaviour, it can be learned with RL and an efficiency reward, the same way tool use was. And it raises an open systems question: if contexts stop being append-only, serving stacks built around prefix caches need something like SCR as a supported feature, with clearly defined approximation behaviour.

The authors name the main risk themselves. A model-editable context is a persistence channel. A prompt injection or a self-generated instruction written into the context survives every later turn, and the paper cites a reported case of a model inserting unauthorised instructions into its own compaction summary. Context integrity now needs its own defences. Two smaller caveats: the headline zero-shot numbers come from strong models, while the 9B model needed RL to beat a summary baseline, and SCR’s stale-KV approximation is validated empirically on these tasks rather than bounded.

Read alongside

  • Zhang, Kraska, Khattab, Recursive Language Models (2025): context as a REPL variable, the closest prior idea, which handles reading but not live-context editing.
  • Li et al., ACM: Agentic Context Management for Long Horizon Tasks (2026), and Li et al., Self-Compacting Language Model Agents (2026): the action-based rung CLMs generalise.
  • Zhou et al., MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents (ICLR 2026): trained consolidation under a harness schedule, one of the main baselines.
  • Shao et al., DeepSeekMath (2024): the GRPO formulation the RL recipe extends stepwise.
  • Chen et al., BrowseComp-Plus (2025): the fixed-corpus deep-research benchmark behind the headline numbers.
  • Sutton, The Bitter Lesson (2019): the argument the authors cite for letting the model search over strategies.

📄 arXiv abstract · 📄 PDF · 💾 facebookresearch/context-language-models


Part of the Weekly CS Paper Digest series. Summary written from a close read of the preprint; figures cropped from the arXiv PDF and reproduced here under fair use for educational commentary.