arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29623cs.CLcs.AI

MI-Distillation:从模型插值的指令推理数据频谱中选择以进行思维链蒸馏

MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation

  • The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yangsong Lan, Renkai Hu, HongKai Zheng, Bo Zhang, Renzhi Wang, Hongliang Dai, Piji Li

AI总结:

本研究提出MI-Distillation框架,通过模型插值构建指令推理数据频谱,引入SeqLSS选择合适轨迹,在推理基准上显著提升了小模型的思维链蒸馏效果。

AI中文摘要:

近期大型推理模型(LRMs)通过长思维链(Long CoT)推理在复杂问题上展现出强性能,但将这类轨迹蒸馏为较小的学生模型仍具挑战性:直接的Long CoT监督常带来有限增益,效果可能不及简洁的短思维链(Short CoT)理由。本研究从以梯度为核心的视角探究该现象,分析发现Long CoT相较于Short CoT会诱导更大的梯度幅度与更集中的更新方向,且该效应随学生模型容量提升而更显著。这些结果表明,有效的Long CoT蒸馏需平衡推理轨迹的推理信息密度及其与学生模型的分布对齐。受此启发,我们提出模型插值蒸馏(Model Interpolation Distillation, MI-Distillation)框架,通过模型插值构建连续的指令推理数据频谱;为从该频谱中选择合适轨迹,进一步引入序列可学习意外度分数(Sequential Learnable Surprisal Score, SeqLSS),其偏好对学生模型兼具信息性与可学习性的推理路径。在推理基准上的大量实验显示,MI-Distillation相比强Long CoT基线,能持续提升小模型的CoT蒸馏效果。

英文摘要:

Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains challenging: direct Long CoT supervision often provides limited gains and can be less effective than concise Short CoT rationales. In this work, we investigate this phenomenon from a gradient-centric perspective. Our analysis shows that Long CoT induces larger gradient magnitudes and more concentrated update directions than Short CoT, with this effect becoming more pronounced as student model capacity increases. These findings suggest that effective Long CoT distillation requires balancing the reasoning information density of reasoning trajectories with their distributional alignment to the student model. Motivated by this insight, we propose \textbf{M}odel \textbf{I}nterporlation \textbf{Distillation} (\textbf{MI-Distillation}), a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation. To select suitable trajectories from this spectrum, we further introduce \textbf{Seq}uential \textbf{L}earnable \textbf{S}urprisal \textbf{S}core (\textbf{SeqLSS}), which favors reasoning paths that are both informative and learnable for the student. Extensive experiments on reasoning benchmarks show that MI-Distillation consistently improves small model CoT distillation over strong Long CoT baselines.

补充信息

↑