arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

特权似然并非自动等价于价值:在线自蒸馏中 token 信用的三项检验

Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation

Xuan-Phi Nguyen, Zeyu Leo Liu, Yang Li, Shrey Pandit, Yiran Zhao, Anurag Koul, Shafiq Joty

arXiv 2608.09263首次发表:更新:

发表机构

Salesforce AI Research(Salesforce AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对在线自蒸馏中 token 信用问题,通过三项检验分析特权似然与价值的关系,实验表明 token 分数变体效果不及仅结果对照组,需单独验证分数等因素的有效性。

AI 中文摘要

结果验证器会对完整的推理轨迹打分,但不会给中间 token 分配信用。特权自蒸馏试图通过仅用训练信息对模型自身的 rollout(推理过程)重新打分来填补这一空白。然而,token 似然变化并不自动等同于结果信用。我们拆分出三个问题:分数是否跟踪更好的动作、反馈构建是否改变了比较对象、训练损失会强化何种行为,我们正式确立了这些区分。当使用针对同一 rollout 编写的事后反馈对其打分时,反馈内容同时决定了 token 和打分上下文,造成直接的自依赖。使用同一问题的另一 rollout 的反馈可消除该依赖,但无法保证分数有用。在 AIME 2025 数据集上针对 20B 模型的匹配实验中,所实现的加性分数接近随机水平(AUC=0.505),且在长度调整后略微偏向错误轨迹。在配对比较中,仅结果的对照组准确率为 64.2%,而五种 token 分数变体的准确率为 24.2%—33.9%。这些结果表明,在将似然信号称为信用之前,需分别验证分数含义、反馈构建和训练行为。

英文摘要

On-policy self-distillation aims to improve upon reinforcement learning from verifiable rewards (RLVR) by providing token-level scores derived from privileged information, such as reference solutions or critic feedback. These scores are treated as estimates of token-level action values, yet they answer a fundamentally different question: how the model's prediction changes when its input context is enriched, rather than how the expected outcome changes when a token is changed. We examine this gap along three dimensions: (i) whether the token-level score tracks task success; (ii) whether feedback generated from the same rollout causes the score to reflect agreement with its own description, and whether using feedback from other rollouts in the group mitigates this self-referential effect; and (iii) what behavior the resulting training objective actually reinforces. In experiments on AIME 2025, the implemented score distinguishes correct from incorrect rollouts at approximately chance level (AUC=0.505); using feedback from a different rollout does not consistently improve this discrimination; and all training configurations achieve only 24.2-33.9% Avg@4, compared with 64.2% for outcome-only GRPO. Moreover, the highest-entropy token decile accounts for 57-71% of the total absolute token-advantage mass, despite the score being least informative about reasoning quality in this regime. By contrast, similar experiments on SciKnowEval Biology improves held-out Avg@8 by 28.0%, while its corresponding trajectory scores achieve AUCs of 0.81-0.92. Together, these results suggest that dense credit assignment through distillation can be effective when its likelihood-based scores are empirically validated as meaningful proxies for outcome-relevant credit. When this alignment does not hold, however, the resulting supervision can fail to generalize and may substantially underperform outcome-based RL.

CommentsUpdated Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑