TRACER:基于结果归因强化学习的LLM智能体单工具上下文保留机制
TRACER: Per-Tool Context Retention for LLM Agents via Consequence-Attributed Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
TRACER将压缩转化为单工具决策问题,通过感知后果的强化学习策略优化上下文保留,大幅降低长程语言智能体的标记消耗且维持任务成功率,具备跨架构迁移性。
中文摘要 AI 辅助
企业数据智能体通过在多个推理步骤中串联大量工具调用来回答业务查询,每个会话通常累积数十万个上下文标记。现有的压缩策略通常在分配保留预算时,未考虑删除单个工具输出的下游后果,因此激进的压缩可能会触发成本高昂的工具重新调用,抵消初始节省,我们将此称为压缩-后果差距。为弥合该差距,我们提出TRACER,它将压缩问题表述为顺序单工具决策问题。轻量级REINFORCE策略仅利用每次压缩事件时可用的信息,分配查询条件下的保留比例。其感知后果的目标函数同时考虑任务成功率、总标记消耗和压缩后的工具重新调用。为改进信用分配,TRACER使用学习到的结果模型,将所选保留比例的预测后果与完全保留每个工具输出的后果进行比较。在三个压缩器后端的保留生产查询上,TRACER与保留所有上下文相比,总标记消耗降低了29%至46%,同时保持相当或更高的任务成功率;与工具类型条件静态策略相比,TRACER额外提供15%至18%的标记节省。干预性部署显示,学习到的单工具信用分数与测得的单工具后果相关;学习到的策略在跨智能体后端和压缩器架构迁移时也能产生正节省,在五个保留的LOCA-bench环境中标记消耗降低了18%至25%。这些结果证明了感知后果的单工具上下文保留对改进长程语言智能体效率的价值。
英文摘要
Enterprise data agents answer business queries by chaining many tool calls over multiple reasoning steps, routinely accumulating hundreds of thousands of context tokens per session. Existing compression strategies typically allocate retention budgets without accounting for the downstream consequences of removing individual tool outputs. Aggressive compression may therefore trigger costly tool re-invocations that offset the initial savings. We call this the compression--consequence gap. To close it, we propose TRACER, which formulates compression as a sequential per-tool decision problem. A lightweight REINFORCE policy assigns query-conditioned retention ratios using only information available at each compression event. Its consequence-aware objective jointly accounts for task success, total token consumption, and post-compression tool re-invocations. To improve credit assignment, TRACER uses a learned outcome model to compare the predicted consequences of the selected retention ratio with those of fully retaining each tool output. On held-out production queries across three compressor backends, TRACER reduces total token consumption by 29--46% relative to keeping all context while maintaining comparable or higher task success. Compared with a tool-type-conditional static policy, TRACER provides an additional 15--18% of token savings. Interventional rollouts show that the learned per-tool credit scores correlate with measured single-tool consequences. The learned policy also yields positive savings when transferred across agent backbones and compressor architectures, and reduces token consumption by 18--25% on five held-out LOCA-bench environments. These results demonstrate the value of consequence-aware, per-tool context retention for improving the efficiency of long-horizon language agents.