Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
过犹不及:强化学习如何减轻大语言模型中任务冲突的综合分析
机构 * Institute of Automation, CAS(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Baidu Inc.(百度公司)
AI总结 研究对比强化学习与监督微调训练的大语言模型的合并行为,经五项任务评估发现强化学习能减少任务冲突及合并后性能下降。通过实验和分析揭示三个关键因素,表明强化学习在模型合并中更具优势。
Comments Published in ICLR 2026