arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35210cs.CL

理解同策略蒸馏:基于稀疏交叉编码器的机制可解释性视角

Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders

  • The University of Hong Kong(香港大学)
  • University of Cambridge(剑桥大学)
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou

AI总结:

通过稀疏交叉编码器和交换读出方法,发现同策略蒸馏不创建或传递特征,而是重新加权学生与教师共享的现有特征,SFT预热提前部分完成此重加权,从而提升蒸馏效果。

AI中文摘要:

同策略蒸馏(On-policy distillation, OPD)是一种被广泛采用的大语言模型推理后训练技术。通常认为它能从更强的教师模型中迁移知识,但OPD实际上将什么蒸馏进学生的内部表征仍不清楚。我们利用稀疏交叉编码器(sparse crosscoders)研究这一问题,该编码器学习一个由OPD前后的学生和教师共享的特征字典。然而,标准的交叉编码器分析能识别模型特有的特征,但无法说明模型如何使用其特征的变化,因为所有模型都被编码到一组特征激活中。因此,我们提出了交换读出(swap readout)方法,该方法独立读取每个学生检查点的特征激活,衡量训练如何改变学生对每个特征的使用,即使对于交叉编码器未见过的检查点也适用。在三种OPD设置中,我们发现OPD既不创建特征,也不传递教师自身的特征,并且学生超过98%的常用特征的触发率变化保持在20%以内。我们进一步研究了通常在OPD之前进行的基于教师轨迹的SFT预热(SFT warm-up),该预热使OPD更有效。预热并非增加特征,而是以两种方式重新加权共享特征。第一,它已经提高和降低了OPD后来会提高和降低的许多特征,提前完成了OPD的部分工作。第二,它改变了OPD单独不会改变的特征,特别是对话格式、推理风格和数学符号相关的特征,并且这些变化在OPD后持续存在。将这种重新加权直接施加于蒸馏学生的特征上,而不改变其权重,使其准确率接近预热学生的水平,而对打乱特征进行相同改变则无此效果。综合这些发现,OPD是重新加权现有特征而非获取新特征:学生从教师那里学习如何使用它们已共享的特征。

英文摘要:

On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify model-specific features but cannot tell how a model's use of its features changes, since all models are encoded into one set of feature activations. We therefore propose the swap readout, which reads each student checkpoint's feature activations on its own, measuring how training changes the student's use of each feature, even for checkpoints unseen by the crosscoder. Across three OPD settings, we find that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over 98% of the student's frequently used features within 20%. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, the warm-up reweights the shared ones in two ways. First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features that OPD alone would not, notably those for conversation format, reasoning style, and mathematical notation, and these changes persist through OPD. Imposing this reweighting on a directly distilled student's features, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD reweights existing features rather than acquiring new ones: the student learns from the teacher how to use the features they already share.

↑