发表机构
Zhejiang University; Shanghai Artificial Intelligence Laboratory, OpenDataLab; Shanghai Jiao Tong University; Tongji University(浙江大学; 上海人工智能实验室,OpenDataLab; 上海交通大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MMVistaReason开放数据后训练方案,通过三组件策略和容量感知训练,在15个多模态基准上以更少样本超越更大模型,实现可靠推理。
AI 中文摘要
开放多模态推理模型受益于大规模推理监督,然而由于数据质量不均、监督构建效率低下、难度分布失衡以及跨领域干扰等问题,可靠的后训练仍然具有挑战性。我们提出了MMVistaReason(MVR),一种开放数据的后训练方案,包含三个组成部分:(1)更广泛的能力覆盖,涵盖互补的分析与真实世界推理组,强调结构化推理与视觉感知和空间定位;(2)高效的SFT和RL数据构建,通过分阶段清洗和标注标准化异构开放数据,结合难度感知的级联教师蒸馏和基于答案似然的轨迹选择来构建MVR-SFT-528K,并应用规模特定的前沿过滤来构建MVR-RL-63K;(3)先专业化后整合的训练,训练互补的RL专家,并通过多教师在线策略蒸馏(MOPD)整合其能力。我们的分析揭示了监督难度、轨迹质量和模型容量之间的容量依赖性交互:较小的学生模型更多受益于精选监督,而较大的学生模型对轨迹变化和混合领域干扰具有鲁棒性。混合领域RL引入了基准级别的负迁移,而MOPD提供了一致的能力整合,且偏好的KL方向随模型规模而变化。在15个多模态基准上,MVR-4B平均得分为72.8,优于Qwen3.5-9B(Instruct)和MMFineReason-8B,同时使用的样本比MMFineReason少约70%。扩展到9B将平均分提升至74.4,超过了Qwen3.5-35B-A3B(Instruct)。总体而言,MMVistaReason证明了系统化的开放数据构建和容量感知的后训练为可靠的多模态推理提供了一条实用且可扩展的路径。
英文摘要
Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.