arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

压缩你所看到的,而非你所说的:面向潜在观测软件工程智能体的锚定上下文蒸馏

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

Zhensheng Zou, Guoqing Wang, Dan Hao

arXiv 2609.31430首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LOHA与ACD方法,通过压缩工具观测为软令牌并蒸馏预测,在减少上下文的同时保持智能体行为,显著提升SWE-bench性能与吞吐量。

AI 中文摘要

工具观测占据了软件工程智能体上下文的主要部分,使得维护长时间交互历史成本高昂。现有的上下文压缩方法可能丢弃后续动作所需的信息,而将智能体适配到软令牌表示则可能损害其原有行为。为了在减少上下文的同时保留动作关键信息和智能体行为,我们结合了潜在观测、硬动作(LOHA)——一种将压缩历史与精确引用所需文本分离的上下文布局——以及锚定上下文蒸馏(ACD)——一种在约束行为漂移的同时实现潜在读取的训练方法。LOHA将较旧的工具观测压缩为软令牌,同时保留智能体自身的轮次和最后K个观测的文本形式,从而提供对历史信息的紧凑访问和对近期内容的精确访问。为使智能体能够使用这种表示,ACD将基础模型的全文预测蒸馏到潜在视图中,同时将其行为锚定在相同基础模型的纯文本输入上。在SWE-bench Verified上,K=3时,Qwen3-4B的每次调用上下文减少43%,SWE-Master-4B-RL减少57%,解决率分别为12.1%和21.8%,而其未压缩基础版本为14.5%和27.5%。单次运行的最近性扫描在K=8时达到14.4%和23.0%,较大的窗口通常更有利于任务性能而非压缩。在32K令牌限制下,Qwen3在K=3时解决了199个实例子集中的21.1%,而相同适配智能体使用全文时仅为11.1%。在并发单GPU服务中,其实例吞吐量达到全文智能体的1.9倍。

英文摘要

Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text needed for exact reference, with Anchored Context Distillation (ACD), a training method that enables latent reading while constraining behavioral drift. LOHA compresses older tool observations into soft tokens while retaining the agent's own turns and the last K observations in text, providing compact access to historical information and exact access to recent content. To enable the agent to use this representation, ACD distills the base model's full-text predictions into the latent view while anchoring its behavior on plain-text inputs to the same base model. On SWE-bench Verified, K=3 reduces context per call by 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with resolve rates of 12.1% and 21.8% versus 14.5% and 27.5% for their uncompressed bases. A single-run recency sweep reaches 14.4% and 23.0% at K=8, with larger windows generally favoring task performance over compression. Under a 32K-token limit, Qwen3 with K=3 resolves 21.1% of a 199-instance subset versus 11.1% for the same adapted agent using full text. In concurrent single-GPU serving, it achieves 1.9 times that full-text agent's instance throughput.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑