arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Latent-MOPD:潜在多教师在线策略蒸馏

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li

arXiv 2610.02381首次发表:更新:

发表机构

Case Western Reserve University; Zillow Group, Inc.; University of Illinois at Chicago; University of Florida(凯斯西储大学; Zillow集团; 伊利诺伊大学芝加哥分校; 佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出首个面向大语言模型的表示级多教师在线策略蒸馏方法,通过整合专家模型的预测与隐藏状态,无需额外教师训练,在数学、代码和逻辑基准上优于多种基线。

AI 中文摘要

在线策略蒸馏(OPD)在智能体生成的响应上训练学生模型。现有的LLM多教师OPD通过输出分布传递专家模型的预测。我们引入了Latent-MOPD,据我们所知,这是首个面向LLM的表示级多教师OPD方法。它通过专家模型的预测以及用于计算这些预测的隐藏状态来整合现有专家,无需额外的教师训练。为了协调来自多个专家的表示监督,我们根据师生关系选择后期层目标,通过共享投影桥接不等的隐藏宽度,并按领域分组更新。每个教师的监督逐渐从隐藏状态转向令牌预测,两个通道使用相同的路由专家。在我们主要的同族设置中,Latent-MOPD在数学、代码和逻辑的九个基准上均优于仅令牌、仅表示和均匀平均基线。在与每个教师参数数量相同的情况下,学生模型在大多数基准上也超过了每个基准的最佳教师。对于更大、独立开发的跨族教师,Latent-MOPD在所有基准上均优于两个单通道基线。同族全层仅表示控制组在领域纯净更新时保持稳定,但在教师领域在更新内交错时崩溃。我们的结果表明,单个学生可以通过输出分布和内部表示整合多个专家的能力。

英文摘要

On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑