arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回归定义:通过轨迹图估计智能体强化学习中的步骤级优势

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang

arXiv 2609.28963首次发表:更新:

发表机构

School of Information Science and Electronic Engineering, Shanghai Jiao Tong University; Tencent AI Platform Department; MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院; 腾讯AI平台部; 上海交通大学人工智能研究院教育部人工智能重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对GRPO等组强化学习方法在步骤级优势估计上的系统性偏差,提出基于轨迹图的忠实步骤级信用分配框架GRAFT,通过图上的贝尔曼迭代恢复状态值并分配信用,并扩展Graph GAE降低偏差,在多轮智能体基准上优于GRPO。

AI 中文摘要

基于组的强化学习(RL)方法,如GRPO及其变体,已成为训练推理和智能体大型语言模型(LLMs)的主要范式。虽然其组归一化的优势估计在响应级别上是可靠的,但在步骤级别上却存在系统性偏差,因为粗粒度的轨迹级优势难以准确反映单个步骤的贡献(即,失败的轨迹可能包含有价值的步骤)。重新审视基础RL定义,我们注意到GRPO在单轮任务上的成功源于其优势估计策略,该策略遵循基本定义:从同一状态采样的多个动作的平均奖励构成可信的状态值估计。将这种忠实估计扩展到步骤级别原则上需要从每个中间状态采样多个动作,这在每个状态基础上代价过高。为缓解此问题,我们提出了一种基于图的忠实步骤级信用分配框架(GRAFT),该框架将所有 rollout 轨迹嫁接成轨迹图,通过图上的贝尔曼迭代恢复节点状态值,并通过节点值差异为每条边分配信用。理论上,估计的步骤级优势忠实遵循RL中的基本优势定义。为进一步确保步骤级优势估计的可靠性,我们进一步提出了Graph GAE,它将GAE扩展到轨迹图以减少状态值估计偏差的影响。在多个多轮智能体基准上的实验显示,与GRPO相比有持续改进,并且性能优于最近的智能体RL算法。代码将在该https URL提供。

英文摘要

Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑