发表机构
Appier AI Research; National Taiwan University(Appier AI Research; 国立台湾大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出$\alpha$Transfer方法,通过在小代理模型上搜索最优合并系数并迁移至大模型,实现高效模型合并,显著提升速度并降低内存消耗。
AI 中文摘要
模型合并通过参数算术将多个微调检查点组合成单个模型,提供了一种有前景的解决方案。然而,寻找最优合并系数需要进行广泛的搜索,随着模型在规模和数量上的扩展,由于高内存需求和搜索空间的组合增长,这种搜索变得代价高昂。我们表明,在同一模型家族内,不同模型大小的模型在合并系数上表现出高度一致的性能分布。这种分布相似性使得一种实用的范式成为可能,我们称之为$\alpha$Transfer:在小代理模型上搜索最优系数,然后直接将其迁移到更大的目标模型。我们在多种合并方法、模型家族和任务上验证了$\alpha$Transfer。实验结果表明,在视觉变换器上实现了6倍加速和70%的内存减少,在大型语言模型上实现了20倍加速和85%的内存减少,同时保持了相当的性能。我们的发现确立了$\alpha$Transfer作为一种高效且可泛化的模型合并扩展方法。
英文摘要
Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{$α$Transfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify $α$Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6$\times$ speedup and 70\% memory reduction on vision transformers, and a 20$\times$ speedup and 85\% memory reduction on large language models, while maintaining comparable performance. Our findings establish $α$Transfer as an efficient and generalizable approach to scaling model merging.
CommentsUnder review