arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

插值策略蒸馏:离线策略与在线策略蒸馏之间的可控连续体

Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation

Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang, Guangting Wang, Fengyun Rao, Jing Lyu, Dong Liu

arXiv 2609.37170首次发表:更新:

发表机构

WeChat Vision, Tencent Inc.; University of Science and Technology of China(腾讯公司微信视觉; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出插值策略蒸馏(IPD),在词元级线性插值学生与教师分布,以可控方式平衡轨迹质量与可学习性,并利用投机解码加速,在文本及多模态推理基准上优于端点策略与启发式方法。

AI 中文摘要

离线策略蒸馏和在线策略蒸馏传统上被视为两种独立的范式,各自偏好蒸馏轨迹的不同属性。教师生成的(离线策略)轨迹通常质量较高,但远离学生的分布;而学生生成的(在线策略)轨迹更易于学习,但往往包含错误的推理。我们将这些范式视为策略连续体的端点,并认为更有效的轨迹策略可能位于两者之间。我们引入了插值策略蒸馏(IPD),它将每个解码步骤的下一词元分布定义为学生分布与教师分布之间的显式线性插值。该插值在分布级别逐词元进行,其系数可直接控制轨迹质量与学生可学习性之间的平衡。直接从此策略采样需要在每个词元处顺序查询教师,因此成本高昂。为使IPD实用化,我们利用一种新的投机解码规则对其进行加速,同时精确保持插值后的下一词元分布。在轨迹层面,由此产生的轨迹自然地在学生生成片段和教师生成片段之间交替。然而,与最近启发式的片段交替方法不同,这种交替是由精确实现的词元级插值策略所诱导的,而非由手工设计的切换规则所决定。在纯文本和多模态推理基准上,IPD始终优于两个端点策略(SFT和OPD)、它们传统的两阶段组合(SFT后接OPD)以及最近的启发式片段交替方法,这表明词元级策略插值能更好地平衡轨迹质量与学生可学习性。

英文摘要

Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbf{Interpolated Policy Distillation (IPD)}, which defines the next-token distribution at every decoding step as an explicit linear interpolation between the student and teacher distributions. The interpolation operates at the distribution level, token by token, and its coefficient provides direct control over the balance between trajectory quality and student learnability. Naively sampling from this policy would require sequentially querying the teacher at every token and is thus expensive. To make IPD practical, we accelerate it with a new speculative-decoding rule while exactly preserving the interpolated next-token distribution.At the trajectory level, the resulting rollouts naturally interleave student- and teacher-generated segments. Unlike recent heuristic segment-interleaving methods, however, this interleaving is induced by an exactly realized token-level interpolated policy rather than by hand-designed switching rules. Across text-only and multimodal reasoning benchmarks, IPD consistently outperforms both endpoint policies (SFT and OPD), their conventional two-stage combination (SFT-then-OPD), and recent heuristic segment-interleaving methods, demonstrating that token-level policy interpolation better balances trajectory quality and student learnability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑