智能体何时可以遗忘其推理?面向长程智能体上下文压缩的ICLR
When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression
浏览论文内容
中文总结 AI 辅助
本研究提出ICLR方法,通过代理熵排序推理块实现长程智能体上下文压缩,在提升任务奖励的同时显著减少令牌消耗,并揭示推理可被安全遗忘的条件。
中文摘要 AI 辅助
长程语言模型智能体会持续累积推理历史,即使早期决策已被执行和观察,上下文长度和推理成本仍会不断增加。与静态的思维链压缩不同,移除历史推理可能会改变未来的动作及由此产生的交互轨迹。我们研究此类推理何时可以被安全地遗忘。我们提出面向长程推理的交互感知压缩方法(ICLR),这是一种无需训练的在线方法,利用冻结的代理熵对推理块进行排序,同时保留动作、工具调用和观察结果。在260个WorkBuddyBench任务上,ICLR将平均奖励从0.699提升至0.718,同时将输入、输出和缓存读取的令牌数分别减少了25.5%、14.4%和33.3%。消融实验揭示了轨迹放大效应,即局部推理删除通过改变后续交互而导致总计算量的非线性变化。表征探测、激活修补和受控轨迹分析进一步表明,一旦与任务相关的派生状态被可靠地外部化到代码、文件、工具输出或环境反馈中,历史推理就变得更加可替代。这些结果将智能体推理表征为动态的工作状态,而非永久的交互历史。
英文摘要
Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR improves average reward from 0.699 to 0.718, while reducing input, output, and cache read tokens by 25.5%, 14.4%, and 33.3%, respectively. Ablations reveal trajectory amplification, where local reasoning deletion produces nonlinear changes in total computation by altering subsequent interaction. Representation probing, activation patching, and controlled trajectory analyses further suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback. These results characterize agent reasoning as dynamic working state rather than permanent interaction history.
发表机构
- TierFlow Team(TierFlow团队)
- Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。