发表机构
University of Southern California; Carnegie Mellon University(南加州大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型在几何问题多模态推理上的不足,构建ODA-Data数据集,提出模态感知互惠推理优化(MIRROR)方法,通过强化学习自我监督,提升多模态推理能力,在相关基准测试中表现更优。
AI 中文摘要
与具有强大推理能力的大语言模型不同,视觉语言模型在视觉推理方面存在困难,即使是在有等效文本、图表和图表+文本视图的几何问题上。不同视图会引发不同行为,标准多模态后训练未充分利用这些互补推理路径和失败模式。为此构建了ODA-Data数据集,开发了模态感知互惠推理优化(MIRROR)方法,通过自我监督改进多模态推理。在几何问题推理基准测试中,MIRROR优于标准强化学习,跨模态行为更准确、一致。
英文摘要
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests that different views expose complementary reasoning paths and failure modes that standard multimodal post-training does not fully exploit. To study and exploit this phenomenon, we construct ODA-Data, a high-quality paired multimodal geometry dataset with text-dominant, image-dominant, and combined image+text views of the same problems, together with splits for training and evaluating modality-dependent reasoning behaviors. We then develop Modality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach for improving multimodal reasoning via self supervision. For each problem, MIRROR evaluates the model under all views, selects the best-performing view as a teacher, and trains other views with a reverse-KL objective towards the teacher. Across reasoning benchmarks that evaluate on geometry problems, MIRROR improves over standard RL and yields more accurate and consistent behavior across modalities