arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过激活引导剪枝实现跨模型规模的无训练知识迁移

Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning

Jiahe Fan, Si Chen, Yinghao Hou, Aiyuan Zhang, Hong Xie

arXiv 2608.13596首次发表:更新:

AI 中文总结

本文提出激活剪枝融合框架APM,通过激活引导选择源模型的显著组件并注入目标模型,无需训练和显式语义对齐,在16个基准上将3B目标模型平均准确率从55.5%提升至60.6%。

AI 中文摘要

异构模型融合旨在结合任务、初始化、架构或规模存在差异的模型。本文研究了一种未被充分探索的跨规模场景:尽管架构存在巨大不匹配,仍用更强的源模型(donor)提升小型目标语言模型(recipient),探究是否无需显式神经元级语义对齐即可迁移有用能力。基于将大模型剪枝至小型架构并注入微小混合权重即可提升目标模型的观察结果,本文提出激活剪枝融合框架(Activation-Prune-Merge, APM)用于跨规模融合。APM在源模型上构建任务条件激活图,选择显著的层、隐藏维度、注意力头和MLP神经元,将其剪枝至目标模型架构,再通过微小插值系数将得到的源模型切片注入原始目标模型。该方法将源模型视为集中功能组件的来源,无需精确结构移植。在涵盖推理、数学、代码生成、指令遵循和分类的16个基准测试中,APM将原始30亿参数目标模型的整体平均准确率从55.5%提升至60.6%;其中RTE准确率从64.3%提升至82.3%,QNLI从52.3%提升至65.7%,BoolQ从70.8%提升至79.2%。对注入比例和顺序多阶段融合的分析进一步表明,激活引导提取可提升可迁移源模型切片的质量,同时保持小比例融合机制。这些结果证明,当源模型贡献足够集中且经过精心选择时,跨规模异构融合无需显式语义对齐即可成功。

英文摘要

Heterogeneous model fusion seeks to combine models that differ in tasks, initializations, architectures, or scales. We study an underexplored cross-scale setting: improving a small recipient language model with a stronger donor despite substantial architectural mismatch. We ask whether useful capabilities can be transferred without explicit neuron-wise semantic alignment. Building on the observation that truncating a large model to a smaller architecture and injecting it with a tiny mixing weight can already improve the recipient, we propose Activation-Prune-Merge (APM), an activation-guided framework for cross-scale fusion. APM constructs task-conditioned activation maps on the donor, selects salient layers, hidden dimensions, attention heads, and MLP neurons to prune it to the recipient architecture, and injects the resulting donor slice into the original recipient using a micro interpolation coefficient. This formulation treats the donor as a source of concentrated functional components rather than requiring precise structural transplantation. Across 16 benchmarks spanning reasoning, mathematics, code generation, instruction following, and classification, APM improves the overall average accuracy from 55.5% to 60.6% over the original 3B recipient. RTE accuracy increases from 64.3% to 82.3%, QNLI from 52.3% to 65.7%, and BoolQ from 70.8% to 79.2%. Analyses of injection ratios and sequential multi-stage fusion further suggest that activation-guided extraction improves the quality of the transferable donor slice while preserving the small-ratio fusion regime. These results provide evidence that cross-scale heterogeneous fusion can succeed without explicit semantic alignment when the donor contribution is sufficiently concentrated and carefully selected.

Comments9 pages, 3 figures, and 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑