PMOPD:多教师在线策略蒸馏中的任务排序、循环与参数更新子空间保护
PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
针对多教师在线策略蒸馏中的能力跷跷板问题,提出基于子空间投影的PMOPD方法,通过任务排序与循环策略减少干扰,在代码、推理和数学任务上显著提升平均得分。
中文摘要 AI 辅助
多教师在线策略蒸馏(MOPD)已成为一种流行的后训练范式,用于在前沿语言模型中整合专门能力。现有的OPD研究主要集中于通过目标设计、蒸馏范围和教师信号构建来优化单任务蒸馏,而MOPD必须在共享参数中聚合多种能力,并解决由此产生的能力跷跷板效应,即提升一个领域会抑制从另一领域获得的能力。受OPD独特更新几何的启发,我们发现MOPD过程中不同任务的参数更新迅速集中在各自低维子空间中,这为识别和控制跨任务干扰提供了直接的几何基础。因此,我们提出PMOPD(基于投影的多教师在线策略蒸馏),该方法从不同任务的累积参数位移构建子空间记忆,并投影梯度和优化器更新,以移除干扰受保护任务方向的成分。我们进一步开发了轻量级冲突探针来表征任务交互并指导任务排序,同时采用循环策略以平衡子空间估计和及时的任务重访。在代表性的代码、推理和数学任务上的实验表明,PMOPD在每项评估能力上均优于MOPD,在Qwen2.5-7B上三项任务的平均得分提高了2.54分,在Llama-3.1-8B上提高了2.09分。这些一致的增益确立了几何感知优化作为实现平衡多教师蒸馏的有效且可迁移的方法。
英文摘要
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
发表机构
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。