arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14171cs.LGcs.CL

分支策略优化:沙盒原生语言智能体强化学习

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

Bowei He, Yankai Chen, Xiaokun Zhang, Xue Liu

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于可执行沙盒的大语言模型智能体强化学习,提出分支策略优化(BPO)算法,利用沙盒特性构建新展开拓扑,通过自适应快照、分叉动作等方式计算优势,实验验证该算法能提升成功率、降低方差并减少策略更新。

中文摘要 AI 辅助

强化学习已成为训练与可执行沙盒交互的大型语言模型智能体的主导范式。诸如PPO、RLOO和GRPO等先进算法从基于人类反馈的强化学习中继承了其展开拓扑结构。本文指出智能体沙盒具有确定性、可快照和可从任何中间状态恢复的特性,这使得可以构建一种不同的展开拓扑结构。在此基础上实例化了分支策略优化(BPO)算法,该算法在骨干轨迹上的高熵决策点自适应地快照沙盒,每个分支点分叉出K个替代动作并展开到终止状态,从兄弟回报计算每步优势。证明了该估计器无偏且方差严格低于轨迹级基线。在WebShop、ALFWorld和SWE-bench上用Qwen2.5 - 7B和Llama-3.1 - 8B骨干模型验证,BPO在匹配计算量时比GRPO和RLOO成功率提高3.6 - 6.1个绝对百分点,梯度范数方差减半,且使用少38%的策略更新就能匹配最佳基线。

英文摘要

Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes. State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topology from RLHF: for each prompt, N independent trajectories are sampled from the initial state, and an advantage is computed by subtracting a group baseline. This design ignores a defining property of agent sandboxes. They are deterministic, snapshottable, and resumable from any intermediate state. We argue that this property enables a fundamentally different rollout topology: rather than N independent trees of depth T, one can construct a single tree of N leaves whose siblings share prefixes, and therefore share variance. We instantiate this idea as Branching Policy Optimization (BPO), a sandbox-native RL algorithm that (i) adaptively snapshots the sandbox at high-entropy decision points along a backbone trajectory, (ii) forks K alternative actions per branch point and rolls out each to termination, and (iii) computes per-step advantages from sibling returns rather than from independent prompts. We prove this estimator is unbiased and has strictly lower variance than the trajectory-level baseline, with the reduction equal to the prefix-explained portion of return variance. On WebShop, ALFWorld, and SWE-bench Verified with Qwen2.5-7B and Llama-3.1-8B backbones, BPO improves success by 3.6--6.1 absolute points over GRPO and RLOO at matched compute, halves gradient-norm variance, and matches the best baseline using 38% fewer policy updates.

发表机构

  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • McGill Univeristy(麦吉尔大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑