arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.15610cs.CLcs.AIcs.LG

用于有效多轮强化学习的过程奖励引导树展开

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

  • UC San Diego(加州大学圣地亚哥分校)
  • Amazon(亚马逊)
  • MIT Alumni(麻省理工学院校友)

机构由 AI 辅助整理,请以论文原文为准。

Xintong Li, Sha Li, Yuwei Zhang, Changlong Yu, Rongmei Lin, Hongye Jin, Shuyi Guan, Xin Liu, Linwei Li, Qingyu Yin, Jingbo Shang

AI总结:

研究多轮强化学习中有效探索问题,提出过程评分器引导的自适应树展开框架PATR,利用过程反馈评分轨迹、选择性分支等,在FrozenLake和SWE - Bench实验中提升性能,是可扩展多轮强化学习的有效策略。

AI中文摘要:

强化学习已成为训练大语言模型智能体的关键方法,然而诸如GRPO/RLOO等流行方法依赖多个独立采样的完整轨迹进行优势估计。在长期代理任务中,这种统一的展开策略会在无信息的死胡同尝试上浪费预算,而有前景的中间状态未得到充分探索。代理轨迹的多轮结构支持将轨迹组组织成树,其中每一轮作为探索的决策点。我们提出了过程评分器引导的自适应树展开(PATR),一种用于多轮代理强化学习的质量感知展开框架。PATR使用任务适当的过程反馈对部分轨迹进行评分,从有前景的状态选择性分支,重用共享前缀,并保守地停止退化路径以减少浪费的采样。实验表明,PATR在FrozenLake和具有挑战性的SWE - Bench上提高了性能,突出了过程引导的树展开作为可扩展多轮强化学习的有效策略。

英文摘要:

Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation. In long-horizon agentic tasks, such a uniform rollout strategy can waste budget on uninformative dead-end attempts, while promising intermediate states do not receive sufficient exploration. The multi-turn structure of agentic trajectories, with interleaved actions and observations, naturally supports organizing a trajectory group as a tree, where each turn serves as a decision point for exploration. This perspective reframes effective exploration as the problem of deciding where to branch. We propose Process-Scorer Guided Adaptive Tree Rollout (PATR), a quality-aware rollout framework for multi-turn agent RL. PATR uses task-appropriate process feedback to score partial trajectories, selectively branches from promising states, reuses shared prefixes, and conservatively stops degenerate paths to reduce wasted sampling. The resulting rollout groups remain compatible with standard policy optimization while providing more efficient exploration under the same training budget. We evaluate PATR on FrozenLake and the challenging SWE-Bench, which is largely unexplored by prior tree-rollout agent RL methods. Experiments show that PATR improves performance by up to +5.0 points on SWE-Bench and +9.3 points on FrozenLake, highlighting process-guided tree rollouts as an effective strategy for scalable multi-turn RL.

补充信息

↑