发表机构
Rensselaer Polytechnic Institute; IBM Research(伦斯勒理工学院; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出激活条件自蒸馏(ACSD),通过对比正确与其余轨迹的激活提取引导向量,在不需参考文本下提升多模型数学推理准确率,最高达71.9%。
AI 中文摘要
在线策略自蒸馏使用模型自身作为教师,为推理提供密集监督,通常通过参考答案条件化实现。提供特权信息本身并不能确保在整个长响应中实现有效的词元级监督。我们提出激活条件自蒸馏(ACSD),该方法通过对比在生成预算内达到验证正确答案的自生成轨迹与所有其余轨迹的激活,提取一个引导向量。基础模型的冻结副本在每个预测位置应用该向量,学生模型则从其在学生生成前缀上的下一词元分布中学习。结果验证用于方向构建和校准;蒸馏既不需要特定问题的参考文本,也不需要教师参数更新。蒸馏后的学生在推理时单独使用。在五个模型中的每一个上,ACSD在四个数学基准上的平均准确率均达到所评估方法中的最高值。在DeepSeek-R1-0528-Qwen3-8B上,平均数学准确率达到71.9%,LiveCodeBench v6 pass@12达到70.9%,而参考条件化的OPSD基线分别为69.0%和66.3%。正确轨迹之间的对比也支持蒸馏,提取的方向可跨数学训练数据集重用。在固定的学生轨迹上,ACSD比OPSD保持更稳定的后期位置对数更新幅度。
英文摘要
On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9\% and LiveCodeBench v6 pass@12 reaches 70.9\%, compared with 69.0\% and 66.3\% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.