arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习蒸馏什么:大语言模型自蒸馏中的双层Top-K令牌选择

Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models

Heng Liang, Xinwen Zhang, Hongchang Gao

arXiv 2610.07247首次发表:更新:

发表机构

Temple University(天普大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有自蒸馏方法无法自适应选择蒸馏令牌的问题,提出基于双层优化的BiToK-SD方法,通过可微阈值松弛实现Top-K令牌动态选择,在数学推理基准上取得最佳平均性能且计算开销小。

AI 中文摘要

大型语言模型展现出强大的推理能力,但其高昂的推理成本使得知识蒸馏成为在资源受限场景下将此类能力迁移至紧凑模型的重要途径。在线策略自蒸馏进一步减少了对外部大型教师模型的依赖,同时提升了紧凑语言模型的推理能力。然而,现有方法通常要么对所有令牌位置进行统一蒸馏,要么使用固定启发式标准选择令牌,对所选位置赋予相同的蒸馏强度,而非自适应地学习哪些令牌对蒸馏最有益。为解决这些局限,我们提出BiToK-SD(自蒸馏的双层Top-K令牌选择),一种基于双层优化的令牌选择方法,用于学习在线策略自蒸馏中应施加蒸馏的位置。具体而言,BiToK-SD被构建为一个双层优化问题,其中下层问题将Top-K令牌选择建模为基于可微阈值的松弛,使所选位置能随学生策略演化而自适应调整,而上层问题则在所选位置执行知识蒸馏。在数学推理基准上的实验表明,BiToK-SD在所有对比方法中取得了最佳平均性能,同时仅需轻量级额外计算。

英文摘要

Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios. On-policy self-distillation further reduces the reliance on external large teacher models while improving the reasoning ability of compact language models. However, existing methods typically either distill all token positions uniformly or select tokens using fixed heuristic criteria, assigning the same distillation strength to the selected positions rather than adaptively learning which tokens are most beneficial for distillation. To address these limitations, we propose BiToK-SD (Bilevel Top-K Token Selection for Self-Distillation), a bilevel-optimization-based token selection method that learns where distillation should be applied during on-policy self-distillation. Specifically, BiToK-SD is formulated as a bilevel optimization problem, where the lower-level problem models Top-K token selection as a differentiable threshold-based relaxation, allowing the selected positions to adapt as the student policy evolves, while the upper-level problem performs knowledge distillation on the selected positions. Experiments on mathematical reasoning benchmarks show that BiToK-SD achieves the best average performance among all compared methods while requiring only lightweight additional computation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑