学习幅度而非方向:面向多轮多步骤大语言模型智能体的验证器约束信用分配
Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents
浏览论文内容
中文总结 AI 辅助
该研究提出CrEST分层信用分配框架,保留RL的验证器约束上限并结合自我教师的密集token级信号,在BFCL V3和WildToolBench上优于基线,可简化教师在策略优化中的作用。
中文摘要 AI 辅助
带可验证奖励的强化学习(RLVR)为训练多轮工具使用智能体提供了验证器约束的性能上限,但其轨迹级信用分配将异构的每轮结果合并为单一奖励信号。在线蒸馏提供了密集的每token监督,但要么受教师约束,要么易出现梯度浓度崩溃。我们提出CrEST,一种分层信用分配框架,既保留了RL的验证器约束上限,又整合了来自特权自我教师的密集token级信号。CrEST在两个层面解决信用分配:按轮段划分的验证优势解决轮间稀释问题,而熵门控的自我教师调制优化轮内token贡献。在BFCL V3和WildToolBench上的实验表明,CrEST在两种模型规模下均始终优于RL和蒸馏基线,在长轨迹和严格会话级指标上的提升最大。本研究证明,教师在策略优化中的作用可从确定更新方向简化为调制更新幅度,在不牺牲验证器约束上限的前提下实现了密集信用分配。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce $\textbf{CrEST}$, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. $\textbf{CrEST}$ resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that $\textbf{CrEST}$ consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.
发表机构
- Zhejiang University(浙江大学)
- Shanghai Innovation Institute(上海创新研究院)
- Westlake University(西湖大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。