arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向智能体强化学习中关键决策的信用分配

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, Andrzej Banburski-Fahey, Jaron Lanier

arXiv 2609.36178首次发表:更新:

发表机构

University of California, Davis; Microsoft; University of Washington; Purdue University(加州大学戴维斯分校; 微软; 华盛顿大学; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ProVer框架,通过智能体评判器识别关键决策片段并验证其优势,实现细粒度信用分配,在多个基准上显著提升GRPO性能。

AI 中文摘要

组相对策略优化(GRPO)已成为训练大型语言模型智能体的有前景方法。然而,其对轨迹级优势的均匀分配未能区分关键决策与次要决策,模糊了哪些中间决策促成了成功。我们提出ProVer,一个在智能体强化学习中针对潜在关键决策进行细粒度信用分配的框架。给定一个回滚组,一个智能体评判器对比成功与失败轨迹,提出一个可能负责其不同结果的片段。ProVer不直接信任评判器的评估,而是通过估计该片段前后采样的当前策略延续之间的终端成功率差异来验证提出的片段。正估计随后被纳入所提片段内策略令牌的GRPO优势中。通过仅使用模型判断来选择验证位置,ProVer将局部信用基于观察到的结果,而无需详尽评估每个中间状态。在ALFWorld、WebShop和SearchQA上,ProVer在两种模型规模下均达到最强平均性能,相对于GRPO,Qwen3.5-2B和Qwen3.5-4B的相对提升分别为9.91%和7.12%。进一步分析表明,即使没有前沿规模的评判器模型,有依据的片段选择也能以适度的额外生成开销改善策略训练,突显了在智能体强化学习中选择性针对关键决策进行细粒度信用分配的有效性和效率。

英文摘要

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑