arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SFT冲突,RL共存:大语言模型多任务学习的理论与实证分析

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Kejian Zhu, Zhuoran Jin, Shangqing Tu, Hongbang Yuan, Yushi Bai, Kang Liu, Juanzi Li, Jun Zhao

arXiv 2608.03573首次发表:更新:

AI 中文总结

该研究针对大语言模型多任务学习,通过理论与实证分析揭示SFT存在任务冲突而RL可稳定共存的机制,进而提出Parallel-RL范式以解耦多任务训练,提升效率与灵活性。

AI 中文摘要

监督微调(SFT)与强化学习(RL)在提升大语言模型(LLM)多任务推理能力方面表现出根本不同的行为。我们的初步实验发现了一种现象:SFT在多阶段训练下会遭受严重的任务冲突,而RL则能让不同任务稳定共存。从实证角度,我们将这一现象追溯到参数层面,观察到RL会在各任务间诱导出稀疏且近似正交的更新。我们通过分析多任务梯度干扰,为该机制提供了理论解释。研究结果揭示了二者的区别:SFT中的干扰受范数限制,与梯度的绝对幅度成比例;而RL中的干扰受方差限制,受优势归一化和在线策略优化所诱导的梯度方差约束。这一较小的方差边界使得各任务间的优化方向近乎正交。利用这一见解,我们提出了Parallel-RL这一范式,它可解耦多任务训练,显著提升了效率与灵活性。

英文摘要

Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the parameter level, observing that RL induces sparse and approximately orthogonal updates across tasks. We provide a theoretical explanation for this mechanism by analyzing multi-task gradient interference. Our results reveal a distinction: interference in SFT is norm-limited, scaling with the absolute gradient magnitude, whereas interference in RL is variance-limited, bounded by the gradient variance induced by advantage normalization and on-policy optimization. This small variance bound yields near-orthogonal optimization directions across tasks. Leveraging this insight, we propose Parallel-RL, a paradigm that decouples multi-task training, significantly improving efficiency and flexibility.

CommentsCode: https://github.com/GaryStack/Parallel-RL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑