arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39369cs.CLcs.LG

探索异构模型合并方法用于复杂知识迁移

Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer

Jiahe Fan, Si Chen, Yinghao Hou, Wenbo Xia, Ke Xu, Hong Xie, Enhong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探索无需训练或对齐的异构参数合并方法,将专用模型能力直接迁移至通用语言模型,实验证明该方法在多种任务中有效。

中文摘要 AI 辅助

专用模型编码了面向任务的行为,但将该行为迁移到通用语言模型通常需要训练、蒸馏或表示对齐。我们研究这种能力是否可以直接在参数层面进行迁移。我们将两种现有的无训练异构合并方法(此前已证明可在通用语言模型之间迁移知识)应用于从专用模型到通用模型的迁移,将专用供体投影到受体的形状,并在无梯度更新或语义对齐的情况下插值骨干参数。交集合并(Intersection-Merge, IM)注入一个与受体形状匹配的前缀对齐供体切片,而激活-剪枝-合并(Activate-Prune-Merge, APM)利用前向传播激活统计来选择在注入前保留哪些供体维度。在嵌入、重排序、奖励建模和混合专家(MoE)代码专用模型迁移中,两种方法均提升了通用受体的性能,表明简单的异构合并可以将能力迁移到多种不同的专用角色中。

英文摘要

Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.

发表机构

  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑