发表机构
Institute of Automation, CAS; University of Chinese Academy of Sciences; Baidu Inc.(中国科学院自动化研究所; 中国科学院大学; 百度公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究对比强化学习与监督微调训练的大语言模型的合并行为,经五项任务评估发现强化学习能减少任务冲突及合并后性能下降。通过实验和分析揭示三个关键因素,表明强化学习在模型合并中更具优势。
AI 中文摘要
模型合并在将多个专业模型整合为一个统一模型中起着关键作用,尤其是在大语言模型时代。近期研究主要聚焦提升合并性能的策略,而训练范式对模型合并有效性的影响未被充分探索。本研究系统探究了强化学习训练的大语言模型与传统监督微调训练的模型的合并行为。通过五项代表性任务的综合评估,发现强化学习显著减少任务冲突且合并后性能下降更少。为揭示原因,进行了大量实证实验和理论分析,发现三个关键因素:强化学习中的策略训练数据控制梯度更新幅度,降低覆盖模型中其他任务现有知识的风险;强化学习优化目标倾向‘过犹不及’,随着模型收敛逐渐减少冲突参数更新的幅度和数量;强化学习中正例和反例的联合优化引导模型趋向无偏的特定任务参数子空间,确保稳健性能并防止参数冲突。
英文摘要
Model merging plays a crucial role in consolidating multiple specialized models into a single, unified model, especially in the era of large language models (LLMs). Recent research has primarily focused on developing strategies to enhance merging performance with the trained models, while the impact of training paradigms, such as supervised fine-tuning (SFT) and reinforcement learning (RL), on the effectiveness of model merging remains underexplored. In this study, we systematically explore the merging behavior of RL-trained LLMs compared to those trained with traditional SFT. Through comprehensive evaluations across five representative tasks, we find that RL significantly reduces task conflicts and results in less performance degradation after merging, making RL-trained models particularly well-suited for this process. To unearth the reasons behind the superior suitability of RL for model merging, we conduct extensive empirical experiments and theoretical analyses. Our findings highlight three key factors: (1) On-policy training data in RL control the gradient updates in a smaller magnitude, reducing the risk of overwriting existing knowledge for other tasks in the model. (2) The RL optimization objective, which favors ``\textit{enough is as good as a feast}", progressively reduces the magnitude and the number of conflict parameter updates as the model converges. (3) Joint optimization of positive and negative examples in RL steers the model towards an unbiased task-specific parameter subspace, ensuring robust performance while further preventing parameter conflicts.
CommentsPublished in ICLR 2026