arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有特权但有偏差:PI条件化教师如何破坏自蒸馏

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation

Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, Akshay Nambi

arXiv 2608.04794首次发表:更新:

发表机构

Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现,以特权信息为条件的自蒸馏在低难度任务上有效,但在困难任务上会因特权信息偏差导致模型推理能力下降,其优化信号与任务成功脱钩。

AI 中文摘要

自蒸馏(Self-distillation, SD)已成为一种计算高效的替代方案,可用于带可验证奖励的强化学习:以关于答案的特权信息(Privileged Information, PI,如参考解决方案)为条件的自教师,为从未见过该信息的学生提供密集的逐词监督。然而,已报道的增益几乎仅来自狭窄、低难度的设置,留下了一个基本问题:作为唯一目标,没有奖励项,SD是否能传授有效知识?我们在其简单设置中复现了SDPO报道的增益,随后将相同设置应用于困难任务,发现SD无法实现增益。在问答、数学、编码和多轮智能体工具使用场景中,跨推理模式、模型规模和PI形式,以及在SDPO和OPSD两种方案下,逐词损失持续下降,但验证准确率未提升,反而通常会下降。我们通过从损失到其产生的模型的单一因果链解释这种失败。该链始于PI偏差:在看到一个特定参考解决方案后,教师的逐词目标被拉向该轨迹,而非普遍正确性,我们用PI偏差分数量化了这种效应。学生被训练为在各处匹配该目标,其目标几乎对rollout是否正确视而不见,其分配的损失主要落在低信息词上,如停用词、标点符号、不确定性标记,而非决定答案的词;在正确的rollout中,探索性词会产生最高差异,因此它会惩罚推理所需的犹豫。结果是一个更平缓、更不果断的学生,其推理能力并未提升:作为唯一目标,SD优化的是与任务成功脱钩的信号。

英文摘要

Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token supervision to a student that never sees it. Reported gains, however, come almost exclusively from narrow, low-difficulty settings, leaving open a basic question: as a lone objective, with no reward term, does SD teach anything? We reproduce SDPO's reported gains in its easy setting, then apply the identical setup to difficult tasks and find that it does not. Across question answering, mathematics, coding, and multi-turn agentic tool use, across reasoning modes, model sizes, and forms of PI, and under both the SDPO and OPSD recipes, the per-token loss falls steadily while validation accuracy does not improve and typically degrades. We explain this failure through a single causal chain from the loss to the model it produces. The chain begins with PI bias: having seen one particular reference solution, the teacher's per-token target is pulled toward that trajectory rather than toward correctness in general, an effect we quantify with a PI Bias Score. Trained to match this target everywhere, the student's objective becomes nearly blind to whether a rollout is correct, and the loss it assigns falls mostly on low-information tokens like stopwords, punctuation, uncertainty markers, rather than those that determine the answer; within correct rollouts the exploratory tokens incur the highest divergence, so it penalizes the hesitation that reasoning requires. The result is a flatter, less decisive student that is no better at reasoning: as a lone objective, SD optimizes a signal decoupled from task success.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑