发表机构
Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Ant Group(中国科学院自动化研究所; 中国科学院大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究从策略蒸馏任务向量的可组合性,发现其能与RL教师更新互补,在任务内和任务间均有效组合,且弱独立性能不意味着弱可组合性。
AI 中文摘要
任务向量通过模型合并提供了一种组合所学能力的简单机制。然而,由从策略蒸馏(OPD)产生的任务向量的可组合性在很大程度上仍未得到探索。OPD利用教师反馈在学生生成的轨迹上训练学生,产生的参数更新与教师模型(通常通过强化学习(RL))产生的更新不同。因此,我们探究OPD任务向量是否能补充其RL教师更新,并在不同任务间有效组合。在五个领域和两种模型架构中,我们找到了两种可组合性的证据。在任务内,合并OPD和RL任务向量可以优于两个组成模型,即使OPD学生弱于其RL教师。在任务间,在八种骨干合并规则比较中,OPD任务向量组合的平均得分高于相应的RL组合。参数空间分析揭示了OPD和RL更新之间存在显著的非共线性。在CODE领域对SMOLLM3-3B的实验表明,在测试的全局更新范数下,组合方向优于任一组成方向,支持该配置下的方向互补性。在任务间,OPD更新在按更新能量排名前10%的前馈通道中也显示出较低的重叠。总之,这些结果表明,较弱的独立性能并不意味着较弱的任务向量可组合性。OPD任务向量可以补充更强的RL教师更新,并在任务间有效组合,凸显了可组合性作为理解和评估训练后更新的一个独特属性。
英文摘要
Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced by on-policy distillation (OPD) remains largely unexplored. OPD trains a student using teacher feedback on student-generated trajectories, yielding parameter updates that differ from those produced by the teacher model, usually by reinforcement learning (RL). We therefore ask whether OPD task vectors can complement their RL teacher updates and compose effectively across tasks. Across five domains and two model architectures, we find evidence for both forms of composability. Within a task, merging OPD and RL task vectors can outperform both constituent models, even when the OPD student is weaker than its RL teacher. Across tasks, OPD task-vector compositions achieve higher average scores than corresponding RL compositions in seven of eight backbone-merging-rule comparisons. Parameter-space analyses reveal substantial non-collinearity between OPD and RL updates. Experiment in CODE domain on SMOLLM3-3B shows that the combined direction outperforms either constituent direction at the tested global update norm, supporting directional complementarity in this configuration. Across tasks, OPD updates also show lower overlap among the top-10% feed-forward channels ranked by update energy. Together, these results show that weaker standalone performance does not imply weaker task-vector composability. OPD task vectors can complement stronger RL teacher updates and combine effectively across tasks, highlighting composability as a distinct property for understanding and evaluating post-training updates.