arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考异构语言模型合并:加权模型平均视角

Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective

Jiahe Fan, Yinghao Hou, Si Chen, Aiyuan Zhang, Hong Xie, Defu Lian

arXiv 2607.18026首次发表:更新:

发表机构

University of Science and Technology of China; School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学; 中国科学技术大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究异构语言模型合并问题,通过无训练的维度适配和比率控制插值,在联合式与交叉式合并中进行实验,结果表明简单参数平均结合轻量级适配和比率控制是强大基线,揭示直接加权融合的限制及对复杂方法的制约。

AI 中文摘要

具有显著不同参数空间的大语言模型能否通过直接加权平均进行合并,而无需训练或语义对齐?现有的异构融合方法通常引入蒸馏、适配器、学习的潜在空间、路由或特征对齐,尚未明确更简单的方法能否适用于真正不同的数十亿参数检查点。我们通过无训练的维度适配和比率控制插值重新审视这个反直觉问题。在联合式合并中,将较小模型扩展到较大参数空间;在交叉式合并中,将较大模型截断到较小参数空间。在跨Qwen系列模型对及涵盖数学推理、代码生成等多种基准测试中,确定性扩展很大程度保留源模型功能,小比率插值可通过转移互补能力改进强源检查点。然而,接近平衡的插值常失败,任务级结果显示存在跷跷板效应。这些结果表明,简单参数平均与轻量级维度适配及精心控制的比率相结合,是异构语言模型合并中出人意料的强大基线,暗示直接加权融合的限制可能也制约更复杂异构合并方法在大规模应用时的效果。

英文摘要

Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment? Existing heterogeneous fusion methods typically introduce distillation, adapters, learned latent spaces, routing, or feature alignment, leaving open whether a simpler recipe can work for genuinely different billion-parameter checkpoints. We revisit this counterintuitive question through training-free dimensional adaptation followed by ratio-controlled interpolation. In union-style merging, we expand the smaller model into the larger parameter space; in intersection-style merging, we truncate the larger model into the smaller parameter space. Across Qwen-family model pairs and benchmarks covering mathematical reasoning, code generation, language understanding, commonsense reasoning, knowledge, and instruction following, deterministic expansion largely preserves the source model function, and small-ratio interpolation can improve over strong source checkpoints by transferring complementary capabilities. However, near-balanced interpolation often collapses, and task-level results reveal a seesaw effect in which gains on some capabilities coexist with regressions on others. These results show that simple parameter averaging, when paired with lightweight dimensional adaptation and carefully controlled ratios, is a surprisingly strong baseline for heterogeneous LLM merging, suggesting that the limits of direct weighted fusion may also bound what more complex heterogeneous merging methods can achieve at scale.

Comments17 pages, 3 figures, 20 tables, preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑