DRSR:学习集合级删除风险以实现高效长时程智能体
DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents
浏览论文内容
中文总结 AI 辅助
针对长时程智能体历史压缩中独立评分忽视集合删除风险的问题,提出DRSR方法,通过反事实监督学习集合级风险并约束选择删除集,在提升奖励的同时显著减少令牌消耗。
中文摘要 AI 辅助
长时程语言模型智能体会累积推理轨迹、工具交互和观察结果,这些内容的相关性会随当前决策而变化。现有的压缩策略通常独立地对历史单元进行评分,但删除多个单元的安全性通常并不由其单独评分决定:冗余证据、累积的微小效应以及删除后剩余的信息都至关重要。我们提出了直接关系集合风险剪枝(DRSR),将智能体历史压缩形式化为对删除集合的风险约束选择。在离线阶段,DRSR通过联合删除协议有效的历史块并测量相同记录的下一个输出的教师强制似然变化来构建精确的反事实监督。然后,一个轻量级评分器根据候选历史与当前行动前状态之间的在线可见关系,以及删除-保留和成对集合结构,来预测集合级危害。在部署时,DRSR使用轻量级评分器评估一小组合法有效的删除候选,并在新近性、协议、预算和学习风险约束下移除最大的可行集合,当没有足够安全的集合时则弃权(不执行)。在WorkBuddyBench Full260上,DRSR将平均奖励从0.699提高到0.802,同时将总模型令牌减少20.820%。在固定的Eval40比较中,它在每个任务1.211M令牌下获得0.794的奖励,比未压缩的智能体少使用35.850%的令牌。机制分析和消融研究进一步表明,决策条件关系、保留上下文信息、成对交互和弃权(不执行)各自对可靠剪枝有所贡献。
英文摘要
Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historical units independently, but the safety of deleting several units is generally not determined by their singleton scores: redundant evidence, accumulated small effects, and the information that remains after deletion all matter. We introduce Direct Relational Set-Risk Pruning (DRSR), which formulates agent-history compression as risk-constrained selection over deletion sets. Offline, DRSR constructs exact counterfactual supervision by jointly deleting protocol-valid history Blocks and measuring the change in teacher-forced likelihood of the same recorded next output. A lightweight scorer then predicts set-level harm from online-visible relations between candidate history and the current pre-action state, together with deleted-retained and pairwise set structure. At deployment, DRSR evaluates a small set of structurally valid deletion candidates with the lightweight scorer and removes the largest feasible set under recency, protocol, budget, and learned-risk constraints, abstaining when no set is sufficiently safe. On WorkBuddyBench Full260, DRSR increases mean reward from 0.699 to 0.802 while reducing total model tokens by 20.820%. On the fixed Eval40 comparison, it obtains 0.794 reward at 1.211M tokens per task, using 35.850% fewer tokens than the uncompressed agent. Mechanistic analyses and ablations further show that decision-conditioned relations, retained-context information, pair interactions, and abstention each contribute to reliable pruning.
发表机构
- TierFlow Team(TierFlow团队)
- Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。