arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习模仿什么:熵感知分布混合

Learning What to Imitate: Entropy-Aware Distribution Mixing

Juan Garcia Giraldo, Matteo Santelmo, Eduard Durech, Imanol Schlag, Valentina Pyatkin, Antoine Bosselut

arXiv 2610.06671首次发表:更新:

发表机构

EPFL; ETH AI Center; Allen Institute for AI(瑞士洛桑联邦理工学院; 苏黎世联邦理工学院人工智能中心; 艾伦人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对小语言模型模仿教师轨迹时的自信冲突,提出熵感知分布混合方法,动态插值学生与教师分布,稳定蒸馏并提升数学推理与泛化能力。

AI 中文摘要

小型语言模型通常在更强的教师模型生成的推理轨迹上进行后训练,作为学生模型以高效学习新技能。然而,在远离学生预期分布的轨迹上进行词元级模仿,往往会产生“自信冲突”,即学生被要求模仿一个它认为不太可能(即低概率)的延续,尽管它对另一个不同的延续很有信心(即处于低熵状态)。为了缓解这些冲突导致的泛化性能下降和灾难性遗忘,我们提出了“熵感知混合”:一种由学生预测熵门控的学生与教师分布的动态逐词元插值方法。我们实现了凸插值和几何插值,用于离线轨迹生成(通过投机解码,然后进行SFT)和在线策略前向KL蒸馏。我们的结果表明,熵感知混合稳定了蒸馏过程,提高了分布内和分布外的数学推理能力,同时比固定教师监督更好地保留了通用能力。尽管如此,最优熵调度取决于训练来源,离线生成的轨迹偏好凹调度(整体教师影响更大),而在线策略训练偏好线性或凸调度(教师集中在高熵状态)。

英文摘要

Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imitate a continuation that it deems unlikely (i.e., low-probability) despite being confident in a different continuation (i.e., in a low-entropy state). To mitigate the degradation in generalisation and catastrophic forgetting caused by these conflicts, we propose \textbf{Entropy-Aware Mixing}: a dynamic per-token interpolation of the student and teacher distributions, gated by the student's predictive entropy. We implement both convex and geometric interpolations for both offline trace generation (via speculative decoding, then SFT) and on-policy forward-KL distillation. Our results show that entropy-aware mixing stabilises distillation, improving in-distribution and out-of-distribution math reasoning while better preserving general capabilities than fixed-teacher supervision. Nonetheless, the optimal entropy schedule depends on the training source, with offline-generated traces favouring concave schedules (greater overall teacher influence) and on-policy training favouring linear or convex schedules (teacher concentrated in high-entropy states).

Comments32 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑