发表机构
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences(中国科学院大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CAMFT提出在微调阶段即考虑可合并性,通过引导任务更新低冲突稀疏坐标,实现高效且兼容的多任务模型合并,实验证明其优于标准微调基线。
AI 中文摘要
模型合并已成为将多种任务特定能力集成到单个大型语言模型中的一种有前景的范式。然而,现有方法主要关注对独立微调模型的事后处理,忽视了训练阶段本身对跨任务兼容性的影响。在微调后解决参数冲突本质上不是最优的。为了解决这个问题,我们提出了CAMFT,一种冲突感知的可合并微调方法,使任务适应既高效又具有合并感知性。CAMFT将可合并性视为在微调过程中形成的属性,而不仅仅是微调后需要解决的问题。通过引导每个任务更新具有较低跨任务冲突的稀疏坐标,CAMFT产生的任务更新既训练高效,又更兼容于下游模型合并。大量实验表明,在多任务合并场景中,CAMFT优于标准微调基线。代码可在该https URL获取。
英文摘要
Model merging has emerged as a promising paradigm for integrating multiple task-specific capabilities into a single large language model. However, existing methods predominantly focus on post-hoc processing of independently fine-tuned models, overlooking how the training phase itself impacts cross-task compatibility. Resolving parameter conflicts after fine-tuning is inherently sub-optimal. To address this, we propose CAMFT, a Conflict-Aware Mergeable Fine-Tuning method that makes task adaptation both efficient and mergeaware. CAMFT treats mergeability as a property shaped during fine-tuning, rather than only a problem to be solved after fine-tuning. By guiding each task to update sparse coordinates with lower cross-task conflict, CAMFT produces task updates that are efficient to train and more compatible for downstream model merging. Extensive experiments demonstrate that CAMFT outperforms standard finetuning baselines in multi-task merging scenarios. Codes are available at https://github.com/gyanchow/CAMFT-LLM.