arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SHARPO:智能体强化学习中的段级信用分配

SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning

Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard

arXiv 2610.00838首次发表:更新:

发表机构

Georgia Institute of Technology; LinkedIn Corporation(佐治亚理工学院; 领英公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SHARPO提出段级信用分配机制,通过师生对数概率差距改进GRPO优势,在ALFWorld和WebShop上超越现有基线。

AI 中文摘要

智能体强化学习(RL)训练大型语言模型(LLM)在长序列、多步骤交互中采取行动。然而,单一的局部错误可能导致任务失败,而轨迹级奖励为单个决策的信用分配提供的指导有限。为解决这一局限性,我们引入了段级后见优势重加权策略优化(SHARPO),一种在环境交互段级别上改进组相对策略优化(GRPO)的信用分配机制。受现有在线策略自蒸馏(OPSD)方法的启发,SHARPO计算每个段内师生对数概率差距,并利用所得信号计算GRPO优势的有界乘数。该乘数由段内所有令牌共享,使得信用在不同段之间可变。使用Qwen2.5-7B-Instruct,SHARPO在ALFWorld和WebShop基准上优于现有基线,包括GRPO、SDAR、RLSD和StepOPSD。

英文摘要

Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.

Comments13 pages, 3 tables, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑