发表机构
University of Pennsylvania; University of Electronic Science and Technology of China; Chongqing University(宾夕法尼亚大学; 电子科技大学; 重庆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉推理中跨模型能力迁移,提出Selective-RL方法,通过选择性迁移RL阶段更新的主导方向,在多个基准上显著提升性能。
AI 中文摘要
模型合并提供了一种无需训练的方式,将推理能力从语言模型迁移到视觉语言模型(VLM),但基于端点的迁移可能会将模型预先存在的差异与推理后训练期间获得的改变混为一谈。我们转而围绕训练阶段的更新来表述能力迁移,隔离出由强化学习(RL)引起的参数变化。然而,完整迁移这一更新仍然不是最优的:我们发现其组成部分在跨模型迁移性上差异显著,其中主导方向比完整更新迁移得更有效。基于这一发现,我们引入了Selective-RL,它隔离RL阶段的更新,保留其主导的矩阵方向并保持幅度,然后将它们迁移到VLM的语言模块。在三个模型家族和五个视觉推理基准上,Selective-RL在12/15的比较中改善了完整更新插值,包括在Qwen接收模型上MathVision提高了8.55个百分点。匹配的对照表明,仅更新幅度或任意低秩并不能重现这些增益。这些结果突显了后训练期间获得的内容与跨模型可迁移内容之间的区别,为跨模型能力迁移提供了训练阶段的视角。代码可在以下网址获取:此https URL。
英文摘要
Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at https://anonymous.4open.science/r/selective-rl.