MetaReason:通过编辑元信息实现精确的交错多模态推理以解决几何问题
MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems
浏览论文内容
中文总结 AI 辅助
本研究提出MetaReason框架,构建TutorGeo与ExamGeo数据集,结合监督微调与强化学习,实现几何问题的精确交错多模态推理,性能优于现有开源模型。
中文摘要 AI 辅助
尽管视觉推理对于解决复杂几何任务至关重要,但现有的视觉语言模型严重依赖仅文本的推理。一些近期方法引入中间视觉状态以促进推理,但它们常受不准确的几何表示和低渲染保真度的阻碍,最终导致不可靠的输出。为解决这些局限,我们提出MetaReason,一个平面几何领域的多模态推理框架,其利用结构化元信息实现精确的辅助线构造。该框架首先将几何图像解析为元信息,使用预定义工具进行可控编辑以合成高保真视觉状态,随后基于这些增强视图开展推理。为支撑该框架,我们构建了TutorGeo,一个包含17000个图像到元信息转换样本、60000个仅文本推理轨迹以及60000个交错多模态推理轨迹的综合数据集。我们结合监督微调与强化学习以开发稳健的多模态推理能力。我们还引入了ExamGeo,一个源自真实考试问题的基准,可实现不同难度级别的系统评估。实验结果表明,MetaReason显著优于现有开源模型,且与专有模型取得了可比的性能。
英文摘要
Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasoning. Some recent methods introduce intermediate visual states to facilitate reasoning, but they are often hindered by inaccurate geometric representations and low rendering fidelity, ultimately leading to unreliable outputs. To address these limitations, we propose MetaReason, a framework for multimodal reasoning in plane geometry that leverages structured meta-information to enable accurate auxiliary-line construction. The framework first parses geometric images into meta-information, performs controllable edits with predefined tools to synthesize high-fidelity visual states, and then conducts reasoning based on these augmented views. To support this framework, we construct TutorGeo, a comprehensive dataset containing 17k image-to-meta conversion samples, 60k text-only reasoning traces, and 60k interleaved multimodal reasoning traces. Using this dataset, we combine supervised fine-tuning and reinforcement learning to develop robust multimodal reasoning capabilities. We also introduce ExamGeo, a benchmark derived from real-world examination problems that enables systematic evaluation across varying difficulty levels. Experimental results demonstrate that MetaReason significantly outperforms existing open-source models and achieves competitive performance against proprietary models.