在模型改变主意之处分叉:用于树结构强化学习的信念转移分支
Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
- University of Minnesota(明尼苏达大学)
- Salesforce AI Research(Salesforce AI 研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出信念转移分支方法,通过检测模型信念变化最大的位置放置分叉,以优化树结构强化学习的步骤级信用分配,在数学和代码任务上显著超越现有基线。
AI中文摘要:
树结构展开为基于可验证奖励的无评论家强化学习(RLVR)提供了步骤级信用分配:在中间点对链进行分叉,兄弟结果差异可估计步骤价值。每次分叉都会增加采样成本,因此实际预算通常只允许每条链进行少量分叉。如果分叉位置设置在结果已基本确定之处,则产生的兄弟结果大多一致,几乎不提供信用信号;因此,在给定树大小的情况下,分叉位置在很大程度上决定了步骤级强化学习能获得的收益。大多数现有主流方法根据结构放置分叉,例如固定长度、中点和分隔符,或根据下一个标记的熵。我们将分叉位置形式化为定位链的价值曲线的“枢轴”,即期望结果发生转折之处。我们提出“信念转移分支”:在候选边界处读取模型的答案信念,并在连续信念分歧最大的步骤之前进行分叉。三种实例化均无需步骤级监督,覆盖不同访问级别:黑盒探针、logit透镜深度剖面以及学习到的激活方向,后者离线拟合,因此仅在强化学习训练前的验证中使用。该信号仅用于“放置”分叉,且探针在数学任务上消耗约步骤计算的1%,在代码任务上低于5%(当它在展开引擎内运行时)。在该验证中,与蒙特卡洛价值曲线相比,信念转移信号在八个模型×基准面板中均排名第一,优于熵、结构和LLM评判基线。在跨三个模型家族和两个领域的强化学习中,信念转移分叉在数学聚合指标上均领先,在OLMo-3-7B上比最强基线高出+2.6聚合值,在AIME 2026上高出+2.9,并横扫所有OLMo代码列,在LiveCodeBench-medium上高出+6.5。
英文摘要:
Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step-level RL can gain. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next-token entropy. We formalize fork placement as locating the \emph{pivots} of the chain's value curve, where the expected outcome turns. We propose \emph{belief-shift branching}: read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most. Three instantiations, none needing step-level supervision, span access levels: a black-box probe, a logit-lens depth profile, and a learned activation direction, which is fit offline and therefore used only in the validation before RL training. The signal only \emph{places} forks, and the probe costs about $1\%$ of step compute on mathematics and under $5\%$ on code when it runs inside the rollout engine. In that validation, against Monte-Carlo value curves, a belief-shift signal ranks first in each of the eight model$\times$benchmark panels, ahead of entropy, structural, and LLM-judge baselines. In RL across three model families and two domains, belief-shift forking leads every mathematics aggregate, on OLMo-3-7B by $+2.6$ aggregate and $+2.9$ on AIME 2026 over the strongest baseline, and sweeps every OLMo code column, by $+6.5$ on LiveCodeBench-medium.