发表机构
SCAI, Arizona State University(亚利桑那州立大学计算与增强智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究能否将大型推理模型中间步骤计算内化到语言模型参数中,提出掩码蒸馏框架,在自我蒸馏和双模型设置中实例化,并改变中间令牌支架长度,通过实验在GSM8K和Countdown推理领域进行评估。
AI 中文摘要
大型推理模型(LRMs)在推理时会产生长的、明确的中间步骤链,这些中间痕迹主导了延迟、内存使用和服务成本,尽管最终答案的正确性与痕迹的正确性没有因果关系,且痕迹长度也不是问题复杂性的可靠指标。因此提出问题:能否将这些中间令牌中的计算内化到语言模型参数中?为此引入掩码蒸馏,这是一种知识蒸馏框架,学生语言模型在问题条件下训练以预测解决方案令牌,推理教师在问题及其思维链痕迹条件下为学生响应提供反馈。在两种设置中实例化该框架:自我蒸馏设置和双模型设置。此外,通过改变监督学生的中间令牌支架的长度,在完全内化(学生只发出解决方案)和不内化(学生在答案前发出完整痕迹)之间进行插值。通过在两个推理领域的控制实验评估该框架。
英文摘要
Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises an obvious question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers with much shorter intermediate traces? We propose masked self-distillation, a knowledge-distillation based post-training framework in which copies of the same model are instantiated as teacher and student, and the student model is trained to internalize all or part of the intermediate trace, thus becoming more efficient at inference. We vary the fraction of intermediate trace the student is trained to internalize, interpolating between full internalization and no internalization. We conduct controlled experiments on two reasoning domains: math and graph coloring. We use the masked self-distillation framework to post-train Qwen3-4B & 8B models. Our results demonstrate that this method can be used to improve task performance while increasing inference efficiency across various domains and model sizes. We systematically analyze whether improved efficiency gain in the post-trained models generalize to OOD problems. We find that masked self-distillation models generalize well for in-domain OOD problems, and the masked self-distillation training does not induce catastrophic forgetting in the student model on out-of-domain problems. Furthermore, our ablation study shows that supervised fine-tuning can train models to produce shorter traces, but at the cost of generalization, highlighting the importance of on-policy training in masked self-distillation.