arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UP-MOPD:多教师在线策略蒸馏中的更新投影

UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation

Taojie Zhu, Jing Jin, Yuan Xia, Chenyang Ding, Qunshan He, Wanke Xia, Tao Sun, Yan Chen, Jian Wang, Jinjie Gu, Tao Feng

arXiv 2610.08398首次发表:更新:

发表机构

Tsinghua University; Ant Group; Zhejiang University(清华大学; 蚂蚁集团; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多教师在线策略蒸馏中梯度冲突导致更新恶化的问题,提出UP-MOPD方法,通过投影违规优化器更新,在医学与通用领域及公共基准上显著提升性能。

AI 中文摘要

从多个教师进行在线策略蒸馏,将不同领域的专业知识整合到单个学生模型中,但冲突的梯度可能阻碍这种整合。在普通SGD下,梯度修正直接约束参数更新。然而,在使用如AdamW等优化器时,动量、自适应缩放和权重衰减可能将修正后的梯度转化为一阶上增加某个领域损失的更新。为解决这一差距,我们提出了多教师在线策略蒸馏的更新投影(UP-MOPD)。UP-MOPD让原始混合梯度更新优化器状态并生成候选位移,然后仅在违反约束的候选提交到参数之前对其进行投影。该投影给出在欧几里得距离上最接近候选的唯一可行更新。在结合医学和通用领域的实验中,UP-MOPD在训练后期将IFEval-loose准确率比普通M-OPD提高了2.96个百分点。它在八个指标上的平均得分为60.03,而梯度投影为59.00,更新拒绝为59.15。在涵盖数学、代码和指令遵循的公共基准上,它在六个任务中取得最佳平均值(32.67),在LiveCodeBench v5上领先,并在IFEval上并列最佳。这些结果支持通过投影优化器更新来减少领域间的干扰。

英文摘要

On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight decay can turn a corrected gradient into an update that increases a domain loss to first order. To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD). UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters. The projection gives the unique feasible update closest to the candidate in Euclidean distance. In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD. It achieves an average score of 60.03 across eight metrics, compared with 59.00 for gradient projection and 59.15 for update rejection. On a public benchmark covering mathematics, code, and instruction following, it achieves the best average across six tasks (32.67), leads on LiveCodeBench v5, and ties for the best IFEval result.These results support projecting optimizer updates to reduce interference between domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑