发表机构
Salesforce AI Research(Salesforce AI研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对在线自蒸馏中 token 信用问题,通过三项检验分析特权似然与价值的关系,实验表明 token 分数变体效果不及仅结果对照组,需单独验证分数等因素的有效性。
AI 中文摘要
结果验证器会对完整的推理轨迹打分,但不会给中间 token 分配信用。特权自蒸馏试图通过仅用训练信息对模型自身的 rollout(推理过程)重新打分来填补这一空白。然而,token 似然变化并不自动等同于结果信用。我们拆分出三个问题:分数是否跟踪更好的动作、反馈构建是否改变了比较对象、训练损失会强化何种行为,我们正式确立了这些区分。当使用针对同一 rollout 编写的事后反馈对其打分时,反馈内容同时决定了 token 和打分上下文,造成直接的自依赖。使用同一问题的另一 rollout 的反馈可消除该依赖,但无法保证分数有用。在 AIME 2025 数据集上针对 20B 模型的匹配实验中,所实现的加性分数接近随机水平(AUC=0.505),且在长度调整后略微偏向错误轨迹。在配对比较中,仅结果的对照组准确率为 64.2%,而五种 token 分数变体的准确率为 24.2%—33.9%。这些结果表明,在将似然信号称为信用之前,需分别验证分数含义、反馈构建和训练行为。
英文摘要
On-policy self-distillation aims to improve upon reinforcement learning from verifiable rewards (RLVR) by providing token-level scores derived from privileged information, such as reference solutions or critic feedback. These scores are treated as estimates of token-level action values, yet they answer a fundamentally different question: how the model's prediction changes when its input context is enriched, rather than how the expected outcome changes when a token is changed. We examine this gap along three dimensions: (i) whether the token-level score tracks task success; (ii) whether feedback generated from the same rollout causes the score to reflect agreement with its own description, and whether using feedback from other rollouts in the group mitigates this self-referential effect; and (iii) what behavior the resulting training objective actually reinforces. In experiments on AIME 2025, the implemented score distinguishes correct from incorrect rollouts at approximately chance level (AUC=0.505); using feedback from a different rollout does not consistently improve this discrimination; and all training configurations achieve only 24.2-33.9% Avg@4, compared with 64.2% for outcome-only GRPO. Moreover, the highest-entropy token decile accounts for 57-71% of the total absolute token-advantage mass, despite the score being least informative about reasoning quality in this regime. By contrast, similar experiments on SciKnowEval Biology improves held-out Avg@8 by 28.0%, while its corresponding trajectory scores achieve AUCs of 0.81-0.92. Together, these results suggest that dense credit assignment through distillation can be effective when its likelihood-based scores are empirically validated as meaningful proxies for outcome-relevant credit. When this alignment does not hold, however, the resulting supervision can fail to generalize and may substantially underperform outcome-based RL.
CommentsUpdated Preprint