Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
机构 * Xi’an Jiaotong University(西安交通大学) ; University of Science and Technology of China(中国科学技术大学) ; SenseTime Research(商汤科技研究院)
Comments Accepted to NeurIPS 2025 (The Thirty-Ninth Annual Conference on Neural Information Processing Systems)