arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22886cs.LGcs.AI

Merge++:通过无数据检查点反演实现通用合并精炼

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Aditya Pola, Vineeth N. Balasubramanian

首次发表
浏览论文内容

中文总结 AI 辅助

Merge++通过反演专家检查点合成任务图像并蒸馏知识,作为无数据后处理阶段,通用提升各类模型合并算法性能,平均增益+2至+8点,最高+25.9点。

中文摘要 AI 辅助

模型合并将微调后的专家模型整合为单个多任务模型,而无需重新训练。所有现有的无数据方法完全在权重空间中处理该问题。这些方法局限于参数上的算术运算,从未观察每个专家模型的行为,而这种行为信号只有通过前向评估才能显现。获取该行为信号需要用于评估的输入,但无数据设置禁止使用额外数据。我们提出Merge++,一种事后方法,通过反演专家检查点来合成任务代表性图像,然后利用这些图像将专家知识蒸馏到合并模型中,从而解决上述问题。Merge++除检查点本身外不需要任何额外数据。它普遍适用于各种合并算法,并作为独立于底层权重空间方法的互补精炼阶段运行。该方法持续改进从简单任务算术到最先进的谱方法等合并算法,平均提升+2至+8个百分点,在个别配置上最高提升+25.9个百分点。

英文摘要

Model merging consolidates fine-tuned experts into one multi-task model without retraining. All existing data-free methods approach this problem entirely in weight space. Restricted to arithmetic on parameters, these methods never observe how each expert behaves, a signal that only emerges through forward evaluation. Accessing this behavioral signal requires inputs to evaluate on, which the data-free setting prohibits. We propose Merge++, a post-hoc method that addresses this by inverting the expert checkpoints to synthesize task-representative images, then distilling expert knowledge into the merged model using those images. Merge++ requires no additional data beyond the checkpoints themselves. It applies universally across merging algorithms and operates as a complementary refinement stage independent of the underlying weight-space method. The method consistently improves merging algorithms ranging from simple task arithmetic to state-of-the-art spectral methods, with average gains of +2 to +8 points and up to +25.9 on individual configurations.

发表机构

  • IIT Hyderabad(印度理工学院海得拉巴分校)
  • Microsoft Research, India(微软研究院(印度))

机构由 AI 辅助整理,请以论文原文为准。

↑