发表机构
University of Virginia; Stanford University(弗吉尼亚大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对同策略自蒸馏抑制探索行为的问题,提出负向自蒸馏框架,通过动态门控机制使模型偏离自生成的缺陷推理,在保留语言能力的同时提升复杂推理性能。
AI 中文摘要
同策略自蒸馏(On-Policy Self-Distillation, OPSD)已成为大语言模型(LLM)自我改进的一种流行范式,它允许模型充当自己的教师,利用诸如真实解答等特权信息。然而,近期研究结果表明,OPSD 在复杂推理任务上会严重降低 LLM 的性能:通过强制学生模型模仿基于特权信息构建的、人为自信的推理轨迹,OPSD 无意中抑制了不确定性的表达,并惩罚了解决难题所需的探索性和自我纠正行为。为解决此问题,我们引入了负向自蒸馏(Negative Self-Distillation, NSD),这是一种新框架,通过偏离有缺陷的推理而非模仿特权解来优化 LLM。NSD 不依赖真实答案或外部监督,而是利用模型自身生成一个针对特定问题的负向条件(例如,扮演一个“粗心推理者”),并推动学生模型的分布远离这个自生成的负向教师。天真地应用去学习(unlearning)目标来实现这种偏离是有问题的,因为有缺陷的推理词元与基本语言词元混杂在一起;不加区分地惩罚两者可能会灾难性地损害模型的基础语言能力。我们通过设计一种动态门控机制来解决此问题,该机制自动识别并隔离推理关键词元,确保梯度更新仅针对行为缺陷,同时保留模型的语言先验。实验上,NSD 持续优于 OPSD 以及其他无标签、自引导的强化学习(RL)基线。
英文摘要
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
Comments23 pages, 7 figures