arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoMerge:面向多任务模型合并的冲突驱动偏好优化

CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging

Mingjie Zheng, Zihao Chen, Wenqing Chen, Weile Yuan, Zhixuan Chu, Jianxing Yu, Zibin Zheng

arXiv 2609.02273首次发表:更新:

发表机构

School of Software Engineering, Sun Yat-sen University; HiThink Research; School of Artificial Intelligence, Sun Yat-sen University; Zhejiang University(中山大学软件工程学院; 海思思考研究院; 中山大学人工智能学院; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CoMerge将模型合并转化为偏好优化问题,利用朴素合并缺陷构建偏好对优化张量级合并系数,在MergeBench和Llama-3.1-8B-Instruct上均优于基线,仅优化少量系数即可实现多任务LLM的高效合并。

AI 中文摘要

模型合并为构建多任务大语言模型(LLM)提供了无需完整模型再训练的高效范式,但仍面临参数干扰问题。现有方法虽旨在保留各专家模型的能力并减轻干扰,但通常未直接从朴素合并(如任务算术)暴露的潜在退化行为中学习。本文提出模型合并的冲突驱动偏好优化框架CoMerge,将模型合并重新表述为偏好优化问题。该方法采用自监督的冲突驱动策略,利用朴素合并方法的缺陷(如任务算术)作为硬负样本,在无需外部标注的情况下构建偏好对。通过应用偏好优化优化轻量级张量级合并系数,CoMerge可减轻参数空间冲突,同时保留任务特定能力。大量实验表明,CoMerge在MergeBench上的平均归一化性能达0.9968,优于所有评估的无数据和有数据模型合并基线。此外,在Llama-3.1-8B-Instruct上,CoMerge在指令遵循、安全等冲突敏感任务上取得显著提升,且仅优化1445个标量系数,仍与全参数微调具有高度竞争力。

英文摘要

Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs) without full model retraining, yet it remains challenged by parameter interference. While existing methods aim to preserve the capabilities of individual expert models and mitigate interference, they generally do not directly learn from the potentially degraded behaviors exposed by naive merging. In this paper, we propose a conflict-driven preference optimization framework for model merging (CoMerge), which reformulates model merging as a preference optimization problem. The approach utilizes a self-supervised, conflict-driven strategy that leverages the defects of naive merging methods (e.g., task arithmetic) as hard negative samples to construct preference pairs without external annotations. By applying preference optimization to refine lightweight, tensor-wise merging coefficients, CoMerge enables the model to mitigate parameter-space conflicts while preserving task-specific capabilities. Extensive experiments show that CoMerge achieves an average normalized performance of 0.9968 on MergeBench, outperforming all evaluated data-free and data-driven model-merging baselines. Furthermore, on Llama-3.1-8B-Instruct, CoMerge yields marked improvements on conflict-sensitive tasks such as instruction following and safety, while remaining highly competitive with full-parameter fine-tuning despite optimizing only 1,445 scalar coefficients.

CommentsAccepted for publication at the EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑