RTPO:用于稳定智能体强化学习训练的反向回合策略优化
RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training
- Adelaide University(阿德莱德大学)
- The University of Sydney(悉尼大学)
- Curtin University(科廷大学)
- School of CSIT(计算机科学与信息技术学院)
- ACFR(澳大利亚空间研究中心)
- School of EECMS(电子工程、计算机与数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多回合智能体RL训练不稳定问题,提出RTPO方法,通过反向回合策略优化消除上下文不匹配与异步漂移,实验显示其较基线提升21.50%、10.76%。
AI中文摘要:
用强化学习(RL)训练多回合智能体工作流,可使大语言模型在单回合设置之外开展复杂推理、使用外部工具及进行迭代搜索。但多回合RL训练仍极不稳定,常随回合数增加出现严重性能下降。通过理论分析,我们确定了三个紧密关联的不稳定来源: rollout-训练的上下文不匹配、稀疏终端奖励下薄弱的回合级信用分配,以及短、长轨迹在不同策略版本下优化时的异步策略漂移。我们表明这些问题在扁平化轨迹优化中具有共同结构根源,并通过统一的反向回合公式解决它们。我们提出反向回合策略优化(RTPO),将多回合rollout组织为稀疏反向树,并按时间反向顺序执行回合级策略更新,使每个决策与其下游延续对齐。RTPO实现因果一致的回合级信用分配和在线策略延续,以控制异步漂移。我们提供理论保证,表明RTPO在提出的回合级公式下消除了上下文不匹配和异步漂移,减少了信用偏差,并收敛到递归最优。在多回合智能体RL基准上的实验显示,RTPO比轨迹级和回合级基线分别提升了21.50%和10.76%,凸显其为使用工具的智能体提供更稳定训练的潜力。
英文摘要:
Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains highly unstable, often causing severe performance degradation as the number of turns increases. Through theoretical analysis, we identify three tightly coupled sources of instability: rollout-training context mismatch, weak turn-level credit assignment under sparse terminal rewards, and asynchronous policy drift when short and long trajectories are optimized under different policy versions. We show that these issues share a common structural origin in flattened trajectory optimization and address them through a unified reverse-turn formulation. We propose Reverse-Turn Policy Optimization (RTPO), which organizes multi-turn rollouts as sparse reverse trees and performs turn-level policy updates in temporal reverse order, aligning each decision with its downstream continuation. RTPO enables causally consistent turn-level credit assignment and on-policy continuation to control asynchronous drift. We provide theoretical guarantees showing that RTPO eliminates context mismatch and asynchronous drift under the proposed turn-level formulation, reduces credit bias, and converges to recursive optimality. Experiments on multi-turn agentic RL benchmarks show that RTPO improves upon trajectory- and turn-level baselines by 21.50% and 10.76%, respectively, highlighting its potential to support more stable training for tool-using agents.