Context Language Models
Context Language Models edit their own context file, beating external managers on agent benchmarks while cutting FLOPs.
Context Language Models treat context as a file the model can update without restriction, including multi-agent setups where several contexts coexist as files. Built zero-shot on existing models, they beat prior context managers with 11.4% higher accuracy and 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores and 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement at the same compute on a 24-hour multi-repository agent swarm. Natural-language skill optimization raised held-out accuracy by up to 35.9 points, online reinforcement learning improved Qwen3.5-9B on BrowseComp-Plus by 47.6% with 12% fewer FLOPs, and Suffix Cache Reuse cut server compute 35% versus standard SGLang.
- CLMs update their own context file instead of relying on an external harness.
- Zero-shot results include 11.4% higher BrowseComp-Plus accuracy and 21.5% fewer FLOPs.
- Online RL lifted Qwen3.5-9B by 47.6% on BrowseComp-Plus with 12% fewer FLOPs.
- Suffix Cache Reuse cut SGLang serving compute by 35% at matched quality.
Full article210 words · extracted from huggingface.co · click to collapse
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.37725