arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨领域整合RLVR能力:融合范式的深入探究

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao

arXiv 2608.27409首次发表:更新:

发表机构

Tencent; Fudan University(腾讯; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究Merge、Mix RL、MOPD三种RLVR融合范式的性能差异与适用场景,发现其平均性能差最多1.4点,单基准差达8.6点,给出了不同场景下的选择指南。

AI 中文摘要

带可验证奖励的强化学习(RLVR)可提升大语言模型的特定能力,但覆盖多种能力通常需要训练独立的领域专家,再对其进行整合。我们根据所复用的人工制品将三种融合范式分类:Merge(合并)结合专家任务向量,Mix RL(混合强化学习)汇集它们的数据集,多教师策略上蒸馏(MOPD)则同时使用前两者。由于这些范式大多被孤立研究,它们的比较方式及选择标准仍不明确。我们在多种模型规模、多领域基准套件中使用共享专家和数据,对三者进行比较。尽管它们的平均性能差异最多为1.4个点,但在单个基准上差距可达8.6个点,且领域层面的变化与任务向量几何中可见的跨领域关系相关。训练动态揭示了不同的约束:Mix RL依赖领域混合比例,MOPD受限于其教师模型,Merge则将所有专家更新压缩为一个。三者均提升了单样本准确率,且未在解决方案覆盖率上取得可测量的增益,也未在保留能力上出现损失。这些结果给出了实用指南:当专家已存在且廉价融合至关重要时使用Merge;当在无专家的情况下训练统一模型,且需调整领域比例以实现跨领域迁移时使用Mix RL;当保留特定领域的增益比超越教师或最小化端到端成本更重要时使用MOPD。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artifacts they reuse: Merge combines expert task vectors, Mix RL pools their datasets, and multi-teacher on-policy distillation (MOPD) uses both. Because they have largely been studied in isolation, how they compare and how to choose among them remain unclear. We compare all three using shared experts and data across model scales and a multi-domain benchmark suite. Although their average performance differs by at most 1.4 points, the gap reaches 8.6 points on a single benchmark, with domain-level variation tracking cross-domain relations visible in task-vector geometry. Training dynamics expose distinct constraints: Mix RL depends on domain mixture proportions, MOPD remains bounded by its teachers, and Merge compresses all expert updates into one. All three improve single-sample accuracy without measurable gains in solution coverage or losses in held-out capabilities. These results yield a practical guideline: use Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters more than surpassing teachers or minimizing end-to-end cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑