arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于推理泛化的几何自蒸馏

Geometric Self-Distillation for Reasoning Generalization

Josip Jukić, Ivan Titov

arXiv 2607.06855首次发表:更新:

发表机构

ILLC, University of Amsterdam; ILCC, University of Edinburgh(阿姆斯特丹大学伊利沙伯格学院; 爱丁堡大学伊利沙伯格中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型特权上下文自蒸馏中监督难以信赖、导致分布外推理能力下降的问题,提出几何自蒸馏目标GeoSD,通过Hellinger损失和近端项对抗漂移,提升了模型分布外推理准确率。

AI 中文摘要

策略内蒸馏是大语言模型实用的训练后方法,为学生模型自身轨迹提供密集教师监督。在特权上下文自蒸馏中,教师和学生是基于相同前缀的同一模型,但教师还能看到提示或完整求解轨迹。这使得监督丰富但难以信赖:教师对特权视角下明显的延续有信心,但学生无法证明其合理性。蒸馏拉力在教师和学生分歧最大处最强,多次更新后会累积成漂移,降低分布外(OOD)推理能力。我们引入了GeoSD,一种几何自蒸馏目标,将这种漂移视为学生预测行为中的移动,并以两种互补方式对抗它。一个Hellinger损失根据学生已有的重叠部分来缩放每个教师偏好,减弱对学生尚无法支持的token的拉力。由于这些拉力在训练中仍会累积,一个近端项会惩罚学生预测与最近检查点的漂移程度,以Fisher-Rao距离衡量。两者都是在下一个token分布的同一几何结构中的距离,自然梯度更新在该几何结构中而不是参数空间中进行步长更新。在数学推理基准和三个模型系列中,GeoSD在保持自蒸馏的分布内收益的同时,相对于基础模型将平均OOD准确率提高了5.7 - 8.6个百分点,且在从1.7B到32B的模型规模上都有提升。分析标准匹配在分布外失败的原因,我们发现它通过从高熵状态的替代方案中抽取质量来赢得与教师的一致,导致对错误答案的自信一致,而GeoSD则使这些替代方案仍可触及。

英文摘要

On-policy distillation provides dense teacher supervision on a language model's own trajectories. In self-distillation with privileged context, this supervision comes from the model itself, conditioned on a hint or solution trace hidden from the student. When the teacher's preferences hinge on privileged information, it can assign higher probability to continuations the student cannot infer from its own context. Matching these preferences throughout training can induce predictive drift and degrade out-of-distribution (OOD) reasoning. We propose GeoSD, a self-distillation method that controls this drift through two complementary geometric terms. A Hellinger loss weights each teacher preference by the student--teacher overlap, reducing the influence of tokens to which the student assigns low probability. Because these influences can still accumulate, a Fisher--Rao penalty regulates predictive distance from a copy of the student refreshed periodically during training. Both terms compare next-token distributions in Fisher--Rao geometry and are jointly optimized with a preconditioner motivated by the natural gradient. Across three model families, GeoSD retains strong in-distribution gains while improving average mathematical OOD accuracy by 5.7--8.6 points over the base model. OOD gains hold across five model scales from 1.7B to 32B and transfer to code generation, where GeoSD improves code accuracy by 1.9 points on average despite distilling on mathematics alone. Our analysis of mathematical reasoning shows that standard matching rapidly concentrates probability mass at high-entropy states and that its samples confidently agree on incorrect answers. In contrast, GeoSD preserves alternative token mass and reduces false consensus.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑