上下文语言模型
Context Language Models
浏览论文内容
中文总结 AI 辅助
本文提出上下文语言模型(CLMs),通过将上下文视为文件并允许模型自由更新,实现原生上下文管理,在多任务上超越现有策略,并支持上下文管理策略的学习与高效服务。
中文摘要 AI 辅助
我们引入了上下文语言模型(CLMs),这是一种能够原生管理自身上下文的语言模型。我们通过将上下文视为一个文件,并允许模型对该文件进行不受限制的更新来实现这一点。这使得模型能够学习上下文中哪些内容最重要并加以维护,并且自然地扩展到多智能体系统,其中多个智能体的上下文作为文件共存。使用现有模型零样本构建的CLMs在多种任务上超越了最先进的上下文管理策略:在BrowseComp-Plus上准确率提高11.4%,FLOPs减少21.5%;在12小时EdgeBench上得分提高5%,FLOPs减少59%;在24小时多仓库智能体群任务中,在相同计算量下改进幅度提高65%。此外,通过将上下文管理从外部框架控制转变为模型内在行为,CLMs自然地实现了上下文管理策略的上下文内学习和参数化学习。我们展示了CLMs可以通过标准技能优化循环演化出的自然语言指令进行引导,在上下文管理任务上将保留准确率提高多达35.9个百分点,同时减少计算量。我们还引入了一种针对CLMs的在线强化学习方法,使Qwen3.5-9B在BrowseComp-Plus上的性能提高47.6%,同时使用减少12%的FLOPs。最后,我们共同设计了用于CLM服务的后缀缓存重用,在匹配性能下相对于标准SGLang进一步减少了35%的服务器端计算量。
英文摘要
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
发表机构
- University of Washington(华盛顿大学)
- Meta Superintelligence Labs(Meta超级智能实验室)
- MIT(麻省理工学院)
- Trillium Labs(Trillium实验室)
机构由 AI 辅助整理,请以论文原文为准。