Context Language Models manage their own context as a file, beating SOTA context management by 11.4% on BrowseComp-Plus with 21.5% fewer FLOPs
- Context Language Models (CLMs) treat the context as a file the model can rewrite without restriction, instead of leaving context management to an external harness; built zero-shot on existing models they beat SOTA context management by 11.4% accuracy on BrowseComp-Plus with 21.5% fewer FLOPs.
- The same zero-shot setup scores 5% higher with 59% fewer FLOPs on 12-hour EdgeBench, and improves 65% more at equal compute on a 24-hour multi-repository agent-swarm task, where multiple agent contexts coexist as files.
- Context-management behavior can be steered with natural-language instructions evolved through a standard skill-optimization loop, raising held-out accuracy by up to 35.9 points on a context-management task while cutting compute.
- An online reinforcement learning method lifts Qwen3.5-9B on BrowseComp-Plus by 47.6% with 12% fewer FLOPs, and Suffix Cache Reuse, co-designed for CLM serving, cuts server-side compute 35% versus standard SGLang at matched performance.
- The paper is by Rulin Shao and 12 co-authors including Luke Zettlemoyer, Mike Lewis, Nathan Lambert, Wen-tau Yih and Pang Wei Koh, submitted to arXiv on 29 September 2026 under cs.AI, cs.CL and cs.LG.
Hacker News opinions
Wow. Context management is one of the big remaining hassles with modern LLMs so this could be big. The obvious complication is cache busting, so it's also exciting they investigated solutions for that.
The biggest discovery might actually be that they ignored regular caching rules, kept invalid cache suffixes, and it didn't hurt performance.
I wonder how bad performance would get if they plain ignored the whole rotary encoding dance and just back-filled precisely the parts of the cache that changed. Would it break the model, confuse it, or would the model internally correct for it?
Can't we do this trick today with any model? Just send the file as next context. Of course you pay for cache misses depending on how deep you make changes, while CLM just ignores the recomputation.
Codex has been moving toward something like this, not released yet. Rather than summary compaction, the model keeps notes as it works and as it approaches the context limit, and a new session is a fresh context with those notes attached plus a pointer back to the previous session. I recreated it in Pi with a max token limit on the note to pressure the model into being concise, and it ends up cheaper than summary compaction too.
Yup, I built a set of fs tools in my custom coding harness that worked exactly like this. You can still get decent caching by ordering things so most of the dynamic stuff comes later.
I'd be concerned about context management eating limited attention. Do you want your agent solving its own memory crisis, or solving the actual task? It can probably do both at once but I suspect there's a non-trivial cost, and a separate hypervisor agent that manages the main agent's context has worked much better for me. You run it on a different schedule and the main agent spends zero tokens thinking about it, plus it's way easier to control when caches get missed.
Maybe have a second model do the management?
So is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come with them?
That was my gut reaction too, but I think it's just a poor abstract. What they actually did looks much deeper than reinventing agents.md.
Eventually the CLM will be a separate model co-trained with the actual model, right? And there will be multiple contexts like hot vs cold pages in a DB. Which reminds me, I'm predicting a 'Context as a DB' paper within one year.
Bitter Lesson showing up yet again, this time in context management?
Related work worth reading alongside this: Recursive Language Models, arXiv 2512.24601.