超越词元对齐:面向跨分词器在线策略蒸馏的事件完成方法
Beyond Token Alignment: Event Completion for Cross-Tokenizer On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
针对跨分词器在线策略蒸馏中教师词元部分生成后的监督缺失问题,提出事件集完成蒸馏(ESCD),通过聚合前缀相关教师事件并监督字节兼容一步完成的总体概率,无需额外采样或修改词表,在数学、代码和科学推理任务上取得一致提升。
中文摘要 AI 辅助
在线策略蒸馏(OPD)通过教师模型对学生生成轨迹的监督,在语言模型之间传递知识。当使用不同的分词器时,一个教师词元可能需要多个学生词元才能生成,从而产生事件已进入但尚未完成的中间状态。现有的跨分词器方法对齐词元或文本片段,以构建可比较的预测目标。我们研究部分生成后的一个互补问题:一旦学生生成了教师词元的前缀,多个后续词元可能完成相同的剩余字节,但教师仅指定所需的完成方式,而非如何在这些有效延续之间分配概率。我们提出事件集完成蒸馏(ESCD),该方法以完成集监督补充跨分词器概率对齐。ESCD聚合与前缀相关的教师事件,并监督字节兼容的一步学生完成的总体概率,从而避免依赖分词器的单个词元间概率分配。该方法复用学生轨迹和预测,既不需要额外的采样,也不改变学生词表。实验表明,在多种模型家族和分词器上,数学、代码和科学推理任务均获得一致提升,并扩展到从1T教师到35B学生的大规模MoE蒸馏。局部分析显示,保留完成集能更好地匹配参考监督,而在所研究的分词器对中,一步完成覆盖了部分事件进入后观察到的超过99%的兼容教师概率质量。这些发现支持将事件进入和事件完成作为跨分词器知识转移的互补监督目标。代码将在GitHub上发布。
英文摘要
On-policy distillation (OPD) transfers knowledge between language models through teacher supervision on student-generated trajectories. With different tokenizers, a single teacher token may require multiple student tokens to generate, creating intermediate states where the event is entered but not yet completed. Existing cross-tokenizer methods align tokens or text spans to construct comparable prediction targets. We study a complementary problem after partial generation: once the student produces a prefix of a teacher token, multiple next tokens may complete the same remaining bytes, but the teacher only specifies the required completion rather than how probability should be divided among these valid continuations. We introduce Event-Set Completion Distillation (ESCD), which complements cross-tokenizer probability alignment with completion-set supervision. ESCD aggregates prefix-related teacher events and supervises the total probability of byte-compatible one-step student completions, avoiding tokenizer-dependent probability splits among individual tokens. The method reuses student trajectories and predictions, requiring neither additional rollouts nor changes to the student vocabulary. Experiments demonstrate consistent gains in mathematics, code, and scientific reasoning across model families and tokenizers, extending to large-scale MoE distillation from a 1T teacher to a 35B student. Local analyses show that retaining completion sets better matches the reference supervision, while one-step completion covers over 99% of observed compatible teacher mass after partial event entry in the studied tokenizer pairs. These findings support event entry and event completion as complementary supervision targets for cross-tokenizer knowledge transfer. Code will be released on GitHub.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。