探索异构模型合并方法用于复杂知识迁移
Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer
浏览论文内容
中文总结 AI 辅助
本研究探索无需训练或对齐的异构参数合并方法,将专用模型能力直接迁移至通用语言模型,实验证明该方法在多种任务中有效。
中文摘要 AI 辅助
专用模型编码了面向任务的行为,但将该行为迁移到通用语言模型通常需要训练、蒸馏或表示对齐。我们研究这种能力是否可以直接在参数层面进行迁移。我们将两种现有的无训练异构合并方法(此前已证明可在通用语言模型之间迁移知识)应用于从专用模型到通用模型的迁移,将专用供体投影到受体的形状,并在无梯度更新或语义对齐的情况下插值骨干参数。交集合并(Intersection-Merge, IM)注入一个与受体形状匹配的前缀对齐供体切片,而激活-剪枝-合并(Activate-Prune-Merge, APM)利用前向传播激活统计来选择在注入前保留哪些供体维度。在嵌入、重排序、奖励建模和混合专家(MoE)代码专用模型迁移中,两种方法均提升了通用受体的性能,表明简单的异构合并可以将能力迁移到多种不同的专用角色中。
英文摘要
Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.
发表机构
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。