arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24870cs.AI

SPO++:面向异步智能体强化学习的流对齐策略优化

SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • Duke University(杜克大学)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang

AI总结:

SPO++通过标准化终端结果优势、按策略事件组织提示证据优化SPO,在ALFWorld和Math-TIR上提升异步智能体强化学习的在线学习效率,动作token测度归一化是其核心有效组件。

AI中文摘要:

组相对强化学习需要等待同一提示的兄弟rollout,对于长且可变的工具使用轨迹而言成本高昂。单流策略优化(SPO)通过持久的提示级价值估计消除了这种依赖,但它在优化token平均actor损失前会对每条轨迹的优势进行白化处理。我们发现轨迹中心化通常不会使actor使用的token加权量中心化,于是通过在动作token测度下标准化终端结果优势来解决这种不匹配。我们还根据生成提示的策略事件而非学习者接收顺序来组织提示证据。在ALFWorld的两个模型规模及Math-TIR的匹配运行中,SPO++相比SPO提升了在线学习效率,配对消融实验表明动作token测度归一化是测试过的最强组件。

英文摘要:

Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level value estimate, but its recipe whitens one advantage per trajectory before optimizing a token-mean actor loss. We show that trajectory centering generally does not center the token-weighted quantity consumed by the actor, and fix the mismatch by standardizing terminal-outcome advantages under the action-token measure. We additionally organize prompt evidence by the policy event that generated it rather than learner receipt order. Across matched runs on ALFWorld at two model scales and on Math-TIR, SPO++ improves online learning efficiency over SPO. A paired ablation identifies action-token-measure normalization as the strongest tested component.

补充信息

↑