发表机构
Shanghai Jiao Tong University; Alibaba Group(上海交通大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出对比策略蒸馏框架COPD,通过冻结教师模型在不同指令下对学生状态评分,以对数概率差为优势信号指导更新,减少推理长度,提升效率,且能集成到自蒸馏框架实现自对比监督。
AI 中文摘要
策略蒸馏(OPD)通过最小化教师和学生在每个token位置的输出分布之间的差异,在从自身策略采样的轨迹上监督学生模型,提供密集的token级监督。现有OPD方法虽能提升学生模型推理能力,但缺乏比较token在不同推理模式下相对兼容性的明确信号。为此提出对比OPD框架COPD,对学生模型生成的每个token,冻结教师模型在两种对比指令下对相同学生状态评分,对数概率之差作为token级优势信号指导OPD更新,直接鼓励学生模型学习更简洁高效的推理策略。在九个多模态基准上实验,结果表明COPD大幅减少推理长度且不影响模型性能,还能无缝集成到策略自蒸馏框架实现自对比监督。
英文摘要
On-policy distillation (OPD) trains a student model on trajectories sampled from its own policy, providing dense token-level supervision by minimizing the divergence between the teacher's and student's output distributions at each token. Although existing OPD approaches effectively distill strong reasoning capabilities into student models, they inherently inherit uncurated chain-of-thought traces, exacerbating overthinking and reasoning redundancy. To address this limitation, we introduce COPD, a contrastive OPD framework. Specifically, for each token generated by the student, a frozen teacher evaluates the current student state under two contrasting prompts that induce low and high reasoning effort, respectively. The resulting difference in log-probabilities serves as a token-level advantage signal to guide the policy update. Rather than merely mimicking a single teacher distribution, COPD guides the student model toward acquiring concise and efficient reasoning strategies. We evaluate COPD across 9 multimodal benchmarks covering both reasoning and understanding tasks. Empirical results show that COPD substantially reduces reasoning length without hurting task performance, consistently improving efficiency across different tasks and model scales. In addition, this contrastive paradigm extends naturally to on-policy self-distillation (OPSD), establishing self-contrastive supervisory signals that enable a single model to compress its own reasoning without an external teacher.
CommentsWork in progress. 33 pages, 12 figures