免训练变压器合并:通过顺序局部算子对齐
Training-Free Transformer Merging via Sequential Local Operator Alignment
- CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
- Technical University of Munich(慕尼黑工业大学)
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出免训练的顺序局部算子对齐方法,沿部分合并模型执行路径合并变压器,通过校准数据顺序对齐算子并分解,减少误差累积,支持秩扩展,跨模态和规模提升多任务合并性能。
AI中文摘要:
免训练模型合并旨在将多个微调模型组合成单个模型,而无需对标记数据进行进一步优化。然而,在变压器中,独立合并各层可能会影响共享的注意力计算,因为查询-键和值-输出算子依赖于组合矩阵,忽略了功能结构。此外,当合并较早的组件时,下游组件接收到的激活与原始模型中的不同,因此合并后的执行路径与原始执行路径不再匹配。在本文中,我们引入了顺序局部算子对齐,这是一种免训练方法,它沿着部分合并模型的执行路径合并变压器。我们的方法使用校准数据来估计每个功能组件的局部行为,在部分合并模型的中间激活下顺序对齐算子,随后将合并后的算子分解回有效的变压器参数。我们经验性地表明,这种顺序步骤减少了跨层的误差累积。此外,所提出的算子分解步骤实现了秩扩展,为增加多任务能力提供了一种有原则的机制。我们证明了我们的方法跨模态、模型规模和不同任务数量具有泛化性,从CLIP和RoBERTa到十亿参数的LLM,并进一步自然地扩展到LoRA微调模型的合并。结果表明,在不需要秩扩展的情况下,我们的方法优于强大的合并基线,而可选的扩展提供了进一步的精度-推理成本权衡。项目链接:此https URL
英文摘要:
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different activations than they do in the original model, thus, the merged and original execution paths no longer match. In this paper, we introduce Sequential Local Operator Alignment, a training-free method that merges transformers along the execution path of the partially merged model. Our method uses calibration data to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and subsequently factorizes the merged operators back into valid transformer parameters. We empirically show that this sequential step reduces error accumulation across layers. Furthermore, the proposed operator factorization step enables rank expansion, providing a principled mechanism for increasing multi-task capacity. We demonstrate that our approach generalizes across modalities, model scales, and varying numbers of tasks, from CLIP and RoBERTa to billion-parameter LLMs, and further extends naturally to the merging of LoRA-fine-tuned models. The results indicate improvements over strong merging baselines without requiring rank expansion, while optional expansion provides a further accuracy-inference-cost trade-off. Project link: https://akansh12.github.io/SLOA-Merge/