arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSPO:用于多轮语言模型智能体的奖励交换策略优化

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents

Qiang Liu, Taian Guo, Ruizhi Qiao, Xing Sun

arXiv 2607.04713首次发表:更新:

发表机构

Tencent YouTu Lab(腾讯优图实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多轮语言模型智能体训练问题,针对稀疏结果奖励训练收敛慢等问题,提出奖励交换策略优化方法,利用密集过程奖励信息,确保轨迹多样性和目标与真实奖励一致性,提升模型性能。

AI 中文摘要

强化学习在训练大型语言模型处理多轮交互任务方面潜力巨大。但在长时多轮任务中,直接用结果奖励训练因信号稀疏和缺乏细粒度反馈收敛慢,且模型可能学不到未采样成功轨迹。定制密集过程奖励虽加速收敛但可能与真实结果奖励不一致。本文提出RSPO,利用密集过程奖励信息促进结果奖励训练,提升模型性能上限。

英文摘要

Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. However, in long-horizon, multi-turn tasks characterized by sparse outcome rewards, directly training with outcome rewards often results in slow convergence due to the sparsity of signals and the lack of fine-grained feedback. Furthermore, the model may fail to learn successful trajectories that are not sampled during training, thereby limiting its performance. Conversely, while employing customized dense process rewards provides richer signals and accelerates convergence, these surrogate rewards may exhibit potential misalignment with the ground-truth outcome rewards. This inconsistency can bias the training direction and ultimately degrade the model's final performance. In this work, we propose Reward-Swap Policy Optimization (RSPO), a method designed to leverage the rich information from dense process rewards to facilitate training with outcome rewards. By utilizing a reward-swap mechanism, RSPO ensures the diversity of sampled trajectories while guaranteeing consistency between the optimization objective and the true outcome rewards, thereby elevating the performance ceiling of the model. We conduct extensive experiments on two challenging agent benchmarks, WebShop and ALFWorld. By applying our method to various reinforcement learning algorithms, including GRPO, PPO, and GiGPO, we demonstrate that RSPO achieves consistent performance improvements across different baselines and benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑