arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向可扩展的RLVR:多模态指令跟随数据合成与蒸馏

Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation

Yirong Zeng, Zhang Sai, Yuxian Wang, Yutai Hou, Yufei Liu, Xiao Ding, Bibo Cai

arXiv 2609.16059首次发表:更新:

发表机构

Harbin Institute of Technology; Peking University(哈尔滨工业大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态指令跟随中RLVR数据稀缺问题,提出MIFS流水线,通过生成式约束协议和可学习性蒸馏合成数据,并利用代码验证器提供奖励,使MLLMs在四个基准上平均提升8.13%,训练收敛快3倍,且缓解了SFT的泛化权衡。

AI 中文摘要

多模态指令跟随(MMIF)对于构建通用智能体至关重要。然而,当前的训练范式严重依赖监督微调(SFT),这往往导致表面层面的模式匹配并降低通用能力。虽然基于可验证奖励的强化学习(RLVR)提供了一种有前景的替代方案,但其在MMIF中的可扩展性受到高质量、可强化学习(RL-ready)多模态数据稀缺的严重制约。为弥合这一差距,我们提出了MIFS(多模态指令跟随合成),一个用于生成可强化学习多模态数据的系统性流水线。具体而言,MIFS引入了一种生成式约束协议来合成多样化的原始样本,随后通过一种基于可学习性的蒸馏机制,根据强化学习训练动态过滤数据,以确保策略优化的稳定性。此外,一个基于代码的验证器为策略学习提供高精度奖励信号。生成的数据集包含8个约束类别和14个任务领域的9万个样本。实证评估表明,使用MIFS训练的多模态大语言模型(MLLMs)在四个MMIF基准上平均提升了8.13%,并且与使用原始数据相比,训练收敛速度提高了3倍。关键在于,我们的方法缓解了SFT典型的泛化权衡,在显著提升指令跟随精度的同时,保留了核心视觉能力。

英文摘要

Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fine-Tuning (SFT), which often leads to surface-level pattern matching and degrades general capabilities. While Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising alternative, its scalability in MMIF is severely bottlenecked by the scarcity of high-quality, RL-ready multimodal data. To bridge this gap, we present MIFS (\textbf{M}ultimodal \textbf{I}nstruction \textbf{F}ollowing \textbf{S}ynthesis), a systematic pipeline designed to generate RL-ready multimodal data. Specifically, MIFS introduces a generative constraint protocol to synthesize diverse raw samples, followed by a learnability-aware distillation mechanism that filters data based on RL training dynamics to ensure stable policy optimization. Furthermore, a code-based verifier provides high-precision reward signals for policy learning. The resulting dataset comprises 90k samples across 8 constraint categories and 14 task domains. Empirical evaluations demonstrate that MIFS-trained MLLMs achieve an average improvement of 8.13\% on four MMIF benchmarks and a 3$\times$ faster training convergence compared to using raw data. Crucially, our approach mitigates the generalization trade-offs typical of SFT, preserving core visual capabilities while significantly boosting instruction-following precision.

Comments14 pages,figures 5

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑