arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13988cs.LG

TRACE:通过信用估计进行长期奖励分配的回合级奖励分配

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多轮智能体训练后的信用分配难题,提出TRACE方法,通过特定状态转换、对数概率获取及转换等步骤进行奖励分配。该方法无需额外训练,在长期复杂搜索任务中显著提升基础模型工具使用能力,在基准测试中表现良好且学习曲线更佳。

中文摘要 AI 辅助

多轮智能体通过一系列工具交互来解决复杂任务,这使得训练后的信用分配成为一个基本挑战。结果奖励对短期推理提供可靠监督,但随着轨迹增长会变得稀疏且方差大,还可能产生误导。我们提出TRACE,一种用于智能体强化学习的密集信用分配方法。TRACE将展开表示为工具调用边界处的状态转换,从冻结的参考模型获取黄金答案对数概率,将其转换为对数比率状态值,并将每个动作的奖励推导为这些值的时间差分变化。这无需额外的评论家或过程标签训练,其单步对数比率TD组件可跨冗余工具调用进行伸缩。在长期复杂搜索中,TRACE通过纯强化学习显著提高了基础模型的工具使用能力,在封闭网络BrowseComp-Plus基准测试中提升了Qwen3-4B和Qwen3-30B-A3B的性能,且学习行为可转移到开放网络基准测试,学习曲线显示在强化学习训练中更早改进和更快收敛。

英文摘要

Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome rewards provide reliable supervision for short-horizon reasoning, but become sparse and high-variance as trajectories grow to tens or hundreds of tool calls. They can also be misleading: a failed rollout may contain many useful actions that move the agent closer to the goal, yet outcome-only training assigns them the same negative advantage as the eventual mistake. We propose TRACE (Turn-level Reward Assignment via Credit Estimation), a dense credit-assignment method for agentic reinforcement learning. TRACE represents rollouts as state transitions at tool-call boundaries, obtains gold-answer log-probabilities from a frozen reference model, transforms them into log-ratio state values, and derives per-action rewards as Temporal-Difference changes in those values. This requires no additional critic or process-label training, and its one-step log-ratio TD component telescopes across redundant tool calls. On long-horizon complex search, TRACE substantially improves base-model tool-use ability using pure RL, without a cold-start supervised fine-tuning stage, an agentic mid-training stage, or training on live-web data. On the closed-web BrowseComp-Plus benchmark, it raises Qwen3-4B from $7.2$ to $35.6$ and Qwen3-30B-A3B from $8.4$ to $42.6$. The learned search behavior also transfers to open-web benchmarks, and the learning curves show earlier improvement and faster convergence during RL training.

发表机构

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑