arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FOCUS:面向LLM智能体的免训练、保持决策的上下文压缩

FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

Shantanu Dixit, Anson Bastos, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan

arXiv 2609.37590首次发表:更新:

发表机构

M365 Research, Microsoft(微软M365研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FOCUS提出免训练的测试时上下文压缩方法,通过因果决策保持重新定义压缩,无需训练即可在多种智能体基准上削减峰值上下文48%、依赖73%,并提升任务成功率8.9个百分点。

AI 中文摘要

LLM智能体积累的交互历史随任务长度线性增长,导致推理成本呈二次方缩放,并因注意力稀释而使性能下降。现有的上下文压缩方法通过离线学习来决定丢弃什么:例如对比优化准则、蒸馏压缩器或训练压缩策略。这会产生大量成本。此外,压缩策略是先验学习的,不能根据测试时轨迹的动态变化进行条件调整。本文提出一个互补性问题:哪些过去的交互在因果上塑造了智能体未来的决策?我们将上下文压缩重新定义为离散交互单元上的因果决策保持问题,并引入FOCUS,一个完全在测试时运行的免训练上下文压缩框架。我们的方法无需离线数据收集或微调,且与架构无关,可作为模块化压缩层附加到任何闭源API前沿模型上。我们在包括API和工具调用、问答、网页领域和多轮对话的多种智能体基准上评估FOCUS。我们的方法建立了新的最先进性能,将峰值上下文削减高达48%,依赖削减73%,同时将任务成功率比未压缩执行提高高达8.9个百分点。

英文摘要

LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.

CommentsPreprint. Under Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑