理解离线与在线策略蒸馏:不同训练目标的故事
Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives
浏览论文内容
中文总结 AI 辅助
本文研究多教师顺序蒸馏中前向与反向KL散度聚合目标的差异,揭示在线策略蒸馏的益处与脆弱性机制,并建立对数遗憾界。
中文摘要 AI 辅助
在线策略蒸馏(OPD)从教师对学生生成响应的反馈中学习,并已在减少相对于监督微调(SFT)的遗忘方面显示出潜力。然而,其益处和脆弱性仍未完全被理解。我们研究从多个教师进行顺序蒸馏,其中学生最小化其与教师的平均散度。前向Kullback-Leibler(KL)散度产生加权算术混合,而反向KL产生归一化加权几何聚合。我们分别开发了在离线策略和在线策略反馈下学习这些目标的算法,在表格设置中建立了对数遗憾界,并将分析扩展到函数逼近。通过分析这些聚合目标,我们识别出有助于解释OPD益处和脆弱性的机制。相对于前向KL,反向KL在无信息反馈下能更好地保留自信专家的偏好,但对教师分配给正确响应极低概率的情况更为敏感。其词元级条件也揭示了对延续分布的依赖,这种依赖可能在长视界内偏向不正确的前缀。
英文摘要
On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。