Shuffle-R1: 通过数据导向的动态洗牌提升多模态大语言模型的强化学习框架
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Huazhong University of Science and Technology(华中科技大学)
- MiLM Plus, Xiaomi Inc.(MiLM Plus,小米公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Shuffle-R1通过动态洗牌和轨迹采样提升多模态大语言模型的强化学习效率,实现更高效的训练效果。
AI中文摘要:
强化学习(RL)已作为一种有效的训练后范式,用于增强多模态大语言模型(MLLM)的推理能力。然而,当前的RL流水线常常由于两个未被充分探索的问题导致训练效率低下:优势坍缩,即批量中的大部分优势集中在零附近,以及回放沉默,即贡献非零梯度的回放比例随时间减少。这些问题导致了次优的梯度更新,并阻碍了长期学习效率。为了解决这些问题,我们提出了Shuffle-R1,一个简单而原则性的框架,通过动态重组轨迹采样和批量组成来提高RL微调效率。它引入了(1)配对轨迹采样,选择具有大优势的高对比度轨迹以提高梯度信号质量,以及(2)基于优势的轨迹洗牌,通过有意识的批量重新排列来增加有价值回放的暴露。在多个推理基准测试中,我们的框架在最小的开销下一致优于强大的RL基线。这些结果突显了数据导向适应对于更高效的MLLM RL训练的重要性。
英文摘要:
Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where most advantages in a batch concentrate near zero, and Rollout Silencing, where the proportion of rollouts contributing non-zero gradients diminishes over time. These issues lead to suboptimal gradient updates and hinder long-term learning efficiency. To address these issues, we propose Shuffle-R1, a simple yet principled framework that improves RL fine-tuning efficiency by dynamically restructuring trajectory sampling and batch composition. It introduces (1) Pairwise Trajectory Sampling, which selects high-contrast trajectories with large advantages to improve gradient signal quality, and (2) Advantage-based Trajectory Shuffle, which increases exposure of valuable rollouts through informed batch reshuffling. Experiments across multiple reasoning benchmarks show that our framework consistently outperforms strong RL baselines with minimal overhead. These results highlight the importance of data-centric adaptations for more efficient RL training in MLLM.