双词汇语言模型用于跨分词器蒸馏
Dual-Vocabulary Language Model for Cross-Tokenizer Distillation
浏览论文内容
中文总结 AI 辅助
提出双词汇语言模型(DVLM),通过并行分词序列和混合前缀注意力解决跨分词器蒸馏中的输入与输出错位,实现分布对齐监督,提升六个推理任务性能。
中文摘要 AI 辅助
同策略蒸馏(OPD)弥合了教师监督与学生行为之间的差距,但不同的师生分词器在输入分词(#1)和输出逻辑(#2)两方面引入了错位。现有方法通过匹配相同文本片段或将分词转换为字节来解决前者,但往往会丢失细粒度的分词信息或破坏原生分词范式;而对于后者,排序、填充或关键分词选择等策略仅保留共享的逻辑维度,导致大量分布信息丢失。在本文中,我们提出了双词汇语言模型(DVLM),该模型用新的学生词汇投影头替换教师的语言模型头,从而获得全维度的学生逻辑(针对#2)。为了支持学生分词(针对#1),它采用并行分词序列(PTS)作为输入,该序列将原始的教师分词序列与一个重新分词序列拼接而成,后者通过将每个学生分词独立转换为教师分词组来形成。为了避免与原始教师分词在推理时的不一致性,混合前缀注意力(HPA)进一步将重新分词组限制在其对应的教师前缀上,并使用其最后状态作为原始学生分词表示的聚合,以投影到学生词汇空间中。类似地,通过PTS和HPA的联合使用,DVLM教师可以在OPD过程中提供与学生输入分词和输出逻辑分布对齐的监督。实验结果表明,我们的DVLM教师具有与原始教师模型相似的收敛损失,并使学生在六个推理任务上的性能得到提升。
英文摘要
On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.
发表机构
- East China Normal University(华东师范大学)
- Shanghai Innovation Institute(上海创新研究院)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。