发表机构
Alibaba Group; Harbin Institute of Technology(阿里巴巴集团; 哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工具集成智能体训练的细粒度信用分配问题,提出CBPO解耦分支采样的预算分配与信用转化,在十个基准上优于现有方法,获两个领域及两种模型规模下最高宏平均准确率。
AI 中文摘要
带可验证奖励的强化学习(RLVR)使语言模型能学习与外部工具的多轮交互,但其稀疏的结果奖励无法提供信号以识别哪些中间决策对成功负责。分支采样在替代延续间引入局部比较,但现有方法常混淆两个不同问题:分配固定的rollout预算,以及将分支结果转化为token级信用。我们提出对比分支策略优化(CBPO),将这两个问题解耦并为每个分配专用机制。生成熵筛选整个响应中的候选分支位置,而路径级和节点级衰减在轨迹和位置间分配固定预算,防止探索坍缩到少数路径或相邻token。父轨迹与共享相同token前缀的分支形成精确前缀组,该受控组内的奖励变化定义了对比分支价值(CBV),这是基于结果的局部决策敏感性估计,可在不改变延续优势符号的情况下重新缩放它们。当沿同一轨迹选择多个节点时,CBPO将其划分为不重叠的信用段,从而避免共享token上的梯度重复。仅需结果奖励而无需过程级注释,CBPO为工具集成智能体训练中的细粒度信用分配提供了实用解决方案。在十个基准上进行的大量实验(包括五个数学推理基准和五个知识密集型搜索基准)表明,CBPO始终优于最先进的策略优化方法和基于分支的方法,在两个领域及两种模型规模下均达到最高的宏平均准确率。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its sparse outcome rewards provide no signal for identifying which intermediate decisions are responsible for success. Branch sampling induces local comparisons among alternative continuations, but existing methods tend to conflate two distinct problems: allocating a fixed rollout budget and translating branch outcomes into token-level credit. We introduce Contrastive Branch Policy Optimization (CBPO), which disentangles these two problems and assigns a dedicated mechanism to each. Generation entropy screens candidate branch positions across the entire response, while path-level and node-level decay distribute a fixed budget across trajectories and positions to prevent exploration from collapsing onto a few paths or adjacent tokens. A parent trajectory together with the branches that share an identical token prefix forms an exact-prefix group, and the reward variation within this controlled group defines the Contrastive Branch Value (CBV), an outcome-based estimate of local decision sensitivity that rescales continuation advantages without altering their sign. When multiple nodes are selected along the same trajectory, CBPO partitions it into non-overlapping credit segments, thereby avoiding duplicated gradients on shared tokens. Requiring only outcome rewards and no process-level annotation, CBPO provides a practical solution for fine-grained credit assignment in tool-integrated agent training. Extensive experiments on ten benchmarks, including five for mathematical reasoning and five for knowledge-intensive search, show that CBPO consistently outperforms state-of-the-art policy-optimization and branch-based methods, attaining the highest macro-average accuracy in both domains and across two model scales.
Comments10 pages, 5 figures, 3 tables