arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19842cs.AI

SAPO:用于智能体强化学习的单回滚自回归策略优化

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Dayang Liang, Lang Feng, Bo An, Yunlong Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对智能体强化学习现有方法的三大局限,提出SAPO框架,其共享策略与价值函数的自回归主干,结合广义优势估计器,在ALFWorld、WebShop任务上性能优于PPO和GRPO,内存与计算效率更高。

中文摘要 AI 辅助

智能体强化学习(RL)已成为大语言模型(LLM)后训练的关键阶段。现有的无价值函数、组相对方法从多轮回滚中估计策略优势,避免了传统近端策略优化(PPO)的大量内存开销,在长视距交互任务上取得了优异性能。尽管这些方法取得了成功,但近期研究揭示了三个局限性:(1)缺乏显式价值泛化和有效的时间信用分配;(2)在长视距复杂任务中存在潜在的优势崩溃问题;(3)需要在采样预算和策略性能之间进行代价高昂的权衡。本研究提出了单回滚自回归策略优化(SAPO),这是一种低内存、计算高效的框架,其中策略和价值函数共享单个自回归主干。SAPO利用LLM的自回归结构,在具有共享参数的不同因果边界处生成策略和价值预测,同时独立优化PPO目标和辅助的在线策略SARSA目标。为了稳健估计每一轮的贡献,我们进一步引入了轨迹级广义优势估计器,将λ回报与批量归一化相结合。在ALFWorld和WebShop上使用Qwen2.5-1.5B/7B开展的实验表明,SAPO训练稳定,性能分别优于PPO和GRPO,平均提升15.1和12.1个百分点,同时消除了单独价值函数模型的内存开销,且每轮迭代的运行时间比PPO减少33.2%。

英文摘要

Agentic reinforcement learning (RL) has emerged as an important post-training approach for enhancing the capabilities of Large Language Models (LLMs). However, existing methods face a trade-off between policy performance and resource efficiency. Conventional Proximal Policy Optimization (PPO) implementations incur substantial memory overhead from a separate critic, whereas critic-free group-relative methods require multiple rollouts and face potential learning bottlenecks on long-horizon tasks. In this work, we propose Single-rollout Autoregressive Policy Optimization (SAPO), an efficient PPO-style framework that unifies policy optimization and value learning within a single causal language model. SAPO exploits the autoregressive structure of LLMs to sequentially generate action and value estimation at distinct causal boundaries with shared parameters, and then jointly optimizes the PPO objectives and an auxiliary on-policy SARSA objective with turn-level generalized advantage estimation, where the latter is designed to facilitate value learning. Extensive experiments on ALFWorld and WebShop with Qwen2.5-1.5B/7B and Qwen3-14B demonstrate that SAPO reduces peak GPU memory usage by 23.1% and per-iteration runtime by 24.8% over strong PPO baseline, while matching or slightly improving task success rate. Our experiments also show that SAPO outperforms Group Relative Policy Optimization (GRPO) and recent cutting-edge variants in both task success and training stability.

发表机构

  • Xiamen University(厦门大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑