选择之前先观察:重新思考在线策略蒸馏中的词汇稀疏化
Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
提出SparseOPD,通过全词汇教师纠正选择重要标记进行稀疏反向传播,在降低内存的同时保持甚至提升在线策略蒸馏性能。
中文摘要 AI 辅助
在线策略蒸馏(OPD)利用教师对学生生成响应的纠正。全词汇纠正即使对于学生分配低概率的标记也能提供重要的纠正,但对所有标记的logits进行反向传播在长序列中会变得内存密集。现有的内存节省方法从采样标记中估计纠正,或将监督限制在学生的TopK标记上,从而引入采样噪声或改变全词汇纠正。我们引入了SparseOPD,它使用全词汇教师纠正来确定哪些纠正重要,然后再选择要微分的标记logits。SparseOPD首先在不保留其反向图的情况下构建全词汇纠正,然后根据纠正幅度而非学生概率选择标记。有符号残差补偿保留了总的促进和抑制纠正质量,而纠正感知的预算分配将稀疏支持分布在各个位置。最后,更新仅通过选定的标记logits进行反向传播。在涵盖数学、化学问答和多模态推理的六个任务-规模设置中,SparseOPD在任务平均准确率上优于采样标记和TopK,并与全词汇相当或超过。在4B数学任务上,梯度余弦相似度达到99%,而8K全参数分析显示反向内存降低70.5%。
英文摘要
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Carnegie Mellon University(卡内基梅隆大学)
- New Jersey Institute of Technology(新泽西理工学院)
- Kuaishou(快手)
机构由 AI 辅助整理,请以论文原文为准。