arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10805cs.CLcs.LG

诊断和缓解基于策略的自蒸馏中的思维崩溃

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan, Liang Ding

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于策略的自蒸馏在复杂推理任务中性能下降问题,发现思维崩溃陷阱。提出自适应双视角OPSD,通过动态调节自蒸馏目标缓解该问题,在多模型规模和数据集上实验验证其有效性,提升绝对平均准确率4.1%。

中文摘要 AI 辅助

基于策略的自蒸馏(OPSD)已成为增强和对齐大语言模型(LLMs)的关键范式。然而,在复杂推理任务中,OPSD反而会降低下游性能。本文系统研究了这一问题,发现了一种严重的优化陷阱——思维崩溃,即模型原生中间推理行为急剧下降。通过基于熵的梯度掩码和令牌级目标分析,揭示了崩溃的触发机制。为解决此问题,提出了自适应双视角OPSD(AD - OPSD),通过不对称逐点散度门动态调节自蒸馏目标。实验表明,AD - OPSD在不同模型规模和数据集上比标准OPSD绝对平均准确率提高了4.1%,能缓解思维崩溃并能稳健推广。

英文摘要

On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In this paper, we systematically investigate this pathology and identify a severe optimization trap we define as \textbf{Thinking Collapse} -- a sharp decline in the model's native intermediate reasoning behavior, measured by epistemic-token density (ET per 1k). Through entropy-based gradient masking and token-level target analysis, we show that this collapse is triggered by aggressive teacher gradients at high-student-entropy decision forks, where student epistemic tokens are frequently suppressed into teacher non-epistemic targets and are highly concentrated in high pointwise student-teacher divergence regions. To resolve this optimization pathology, we propose \textbf{Adaptive Dual-Perspective OPSD (AD-OPSD)}, a robust control framework that dynamically moderates the self-distillation objective. AD-OPSD selectively anchors high-suppression-risk sandboxed tokens to a reference prior derived from the frozen base model via an asymmetrical pointwise divergence gate, preserving native thinking capacity while retaining OPSD's error-correcting power. Extensive experiments across competitive mathematical benchmarks show that AD-OPSD improves over standard OPSD by up to \textbf{+4.1\%} absolute average accuracy across diverse model scales and datasets. Further analysis demonstrates that AD-OPSD mitigates thinking collapse and generalizes robustly to different post-training paradigms.

发表机构

  • Beihang University(北京航空航天大学)
  • Hong Kong Polytechnic University(香港理工大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑