揭开在线策略蒸馏的神秘面纱:作用、问题及调控
Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
浏览论文内容
中文总结 AI 辅助
研究在线策略蒸馏的作用、问题及调控,阐明其为探索催化剂,揭示师生不匹配和长度利用问题,提出优势裁剪和对数尺度压缩调控,实验表明良好调控的信号质量决定OPD中成功探索。
中文摘要 AI 辅助
在线策略蒸馏(OPD)已成为大语言模型训练后的关键范式,但其训练动态仍未被充分理解。我们进行了一项系统研究,考察OPD的作用、问题及调控。首先阐明OPD作为探索催化剂的作用:通过密集的token级指导引导学生走向正确推理路径,而不提高能力上限。通过表明提示多样性比每个问题的采样数量更重要,且OPD的有效性完全取决于其指导信号的质量来证实这一点。这种依赖性揭示了两种阻碍探索的问题。当师生分布差距大导致指导信号与任务正确性不一致时,会出现师生不匹配,引导探索走向适得其反的方向。当聚合的token级目标产生长度依赖的捷径时,会出现长度利用问题,使学生通过响应截断或冗余填充来操纵奖励格局,探索退化的长度模式而非推理策略。为解决这些问题,我们研究了轻量级信号调控:优势裁剪和对数尺度压缩,确保探索由可靠信号引导。在七个基准上的实验表明,这些调控减轻了长度利用问题并实现了有效蒸馏,稳定超越OPD变体和RLVR基线,从而证实良好调控的信号质量而非仅仅教师规模决定了OPD中成功的探索。
英文摘要
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Tencent AI Lab(腾讯人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。