发表机构
University of Chinese Academy of Sciences; KwaiKAT Team; Zhejiang University(中国科学院大学; 快手KwaiKAT团队; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何通过策略内蒸馏整合不同家族语言模型,提出字节前缀边缘化方法,在共享字节空间重表达教师下一个令牌分布,在数学和编程基准测试中,该方法优于现有跨分词器方法,提升了平均准确率。
AI 中文摘要
来自不同家族的开放权重语言模型展现出互补能力,促使通过策略内蒸馏(OPD)将它们整合为一个紧凑的学生模型。然而,全词汇 OPD 通常假定有共享分词器,而现有跨分词器方法可能会丢弃教师概率质量或把它分配给内容不相关的学生令牌。我们引入字节前缀边缘化(BPM),它在共享字节空间中重新表达教师在学生词汇表上的下一个令牌分布。具体而言,BPM 将每个教师令牌的概率分配给字节表示是教师令牌字节前缀的最长学生令牌……在六个数学和编程基准测试中,BPM 始终优于当前跨分词器方法,比最强基线在六个基准测试的平均准确率上提高了 3.7 - 6.6 分。
英文摘要
Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, while existing cross-tokenizer methods may discard teacher probability mass or assign it to student tokens with unrelated content. We introduce Byte-Prefix Marginalization (BPM), which re-expresses the teacher's next-token distribution over the student vocabulary in a shared byte space. Specifically, BPM assigns each teacher token's probability to the longest student token whose byte representation is a prefix of the teacher token's bytes, aggregates mass mapped to the same student token, and places otherwise unmatched mass in an explicit residual category. This produces a vocabulary-complete, byte-aligned, and mass-preserving target for dense OPD. The target exactly recovers the teacher-induced byte-prefix marginal when the relevant prefix does not span multiple teacher tokens (a condition satisfied at more than 99% of training positions) and uses a mass-preserving, chain-factorized lower bound otherwise. Across Qwen3-32B, GLM-Z1-9B-0414, and MiniMax-M2.7 as teachers, BPM consistently outperforms current cross-tokenizer methods on six mathematics and programming benchmarks, improving six-benchmark avg@8 by 3.7-6.6 points over the strongest baselines.
CommentsProject page: https://bpm-opd.github.io/