跨回放贝尔曼闭包用于长视界智能体强化学习
Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
提出跨回放贝尔曼闭包(CRBC),通过合并回放组为经验过程并求解行为策略贝尔曼不动点,在共享锚点间传播证据以提供步骤级信用,在多个基准上显著提升长视界智能体强化学习的性能与效率。
中文摘要 AI 辅助
基于组的强化学习方法(如GRPO)通过比较每个任务采样的回放来训练LLM智能体,而无需学习评论家。在长视界设置中,这些回放会重新访问共享锚点状态,为步骤级信用提供跨回放证据。理想情况下,步骤级信用应纳入锚点处观察到的已实现后缀之外的证据,同时根据经验频率聚合替代续段。访问局部平均在共享锚点处汇集已实现的后缀回报并尊重观察到的频率,但不会跨回放递归传播证据,而最短路径估计器具有全局可达性,但允许罕见观察到的路线主导锚点的价值。我们引入了跨回放贝尔曼闭包(CRBC),它将每个回放组合并为一个具有吸收成功和失败边界的有限经验过程,并通过一次线性求解评估其行为策略贝尔曼不动点。该不动点使用相同的经验动作和转移频率,通过共享锚点传播证据并聚合替代续段。通过观察到的转移备份所得的状态值产生动作值,其相对于相应状态值的增益提供步骤级信用。相应的有限深度族在零深度时恢复访问局部回报平均,并随深度增加收敛到精确闭包。归一化的闭包信用与轨迹级组优势相结合用于策略优化,无需额外的环境回放。在ALFWorld、WebShop和Sokoban基准测试中,使用多种模型规模,CRBC持续提高最终性能和学习效率。例如,在ALFWorld上使用Qwen2.5-1.5B-Instruct,CRBC比最强评估基线高出5.59个百分点。
英文摘要
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence beyond the realized suffixes observed at an anchor while aggregating alternative continuations according to their empirical frequencies. Visit-local averaging pools realized suffix returns at shared anchors and respects observed frequencies, but does not recursively propagate evidence across rollouts, whereas shortest-path estimators have global reach but allow a rarely observed route to dominate an anchor's value. We introduce Cross-Rollout Bellman Closure (CRBC), which merges each rollout group into a finite empirical process with absorbing success and failure boundaries and evaluates its behavior-policy Bellman fixed point with one linear solve. This fixed point uses the same empirical action and transition frequencies to propagate evidence through shared anchors and aggregate alternative continuations. Backing up the resulting state values through observed transitions yields action values, whose gain over the corresponding state value provides step-level credit. A corresponding finite-depth family recovers visit-local return averaging at zero depth and converges to the exact closure as depth increases. The normalized closure credit is combined with the trajectory-level group advantage for policy optimization, without additional environment rollouts. Across ALFWorld, WebShop, and Sokoban benchmarks with multiple model scales, CRBC consistently improves final performance and learning efficiency. For example, CRBC outperforms the strongest evaluated baseline by 5.59 percentage points on ALFWorld with Qwen2.5-1.5B-Instruct.
发表机构
- Beihang University(北京航空航天大学)
- Zhongguancun Academy(中关村学院)
- Communication University of China(中国传媒大学)
- Peking University(北京大学)
- Hangzhou Innovation Institute of Beihang University(北京航空航天大学杭州创新研究院)
机构由 AI 辅助整理,请以论文原文为准。