arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34386cs.CLcs.LG

选择之前先观察:重新思考在线策略蒸馏中的词汇稀疏化

Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation

Yongliang Miao, Shuang Liu, Yanguang Liu, Yandong Bai, Mengnan Du

首次发表
浏览论文内容

中文总结 AI 辅助

提出SparseOPD,通过全词汇教师纠正选择重要标记进行稀疏反向传播,在降低内存的同时保持甚至提升在线策略蒸馏性能。

中文摘要 AI 辅助

在线策略蒸馏(OPD)利用教师对学生生成响应的纠正。全词汇纠正即使对于学生分配低概率的标记也能提供重要的纠正,但对所有标记的logits进行反向传播在长序列中会变得内存密集。现有的内存节省方法从采样标记中估计纠正,或将监督限制在学生的TopK标记上,从而引入采样噪声或改变全词汇纠正。我们引入了SparseOPD,它使用全词汇教师纠正来确定哪些纠正重要,然后再选择要微分的标记logits。SparseOPD首先在不保留其反向图的情况下构建全词汇纠正,然后根据纠正幅度而非学生概率选择标记。有符号残差补偿保留了总的促进和抑制纠正质量,而纠正感知的预算分配将稀疏支持分布在各个位置。最后,更新仅通过选定的标记logits进行反向传播。在涵盖数学、化学问答和多模态推理的六个任务-规模设置中,SparseOPD在任务平均准确率上优于采样标记和TopK,并与全词汇相当或超过。在4B数学任务上,梯度余弦相似度达到99%,而8K全参数分析显示反向内存降低70.5%。

英文摘要

On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Carnegie Mellon University(卡内基梅隆大学)
  • New Jersey Institute of Technology(新泽西理工学院)
  • Kuaishou(快手)

机构由 AI 辅助整理,请以论文原文为准。

↑