发表机构
Zhejiang University; Tencent; HKUST (GZ); Zhongguancun Academy; VNET Group; Zhejiang Lab(浙江大学; 腾讯; 香港科技大学(广州); 中关村学院; 世纪互联; 之江实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大模型在数学函数推理中忽视视觉线索的模态干扰问题,提出Func-R1,通过解耦架构、分层后训练和感知对齐理论优化策略,在开源模型中达到最优,并在MathVerse上超越GPT-5达8.4%。
AI 中文摘要
在视觉情境中进行深思熟虑的数学推理是先进多模态大语言模型(MLLMs)的标志性能力,需要感知基础与符号逻辑的复杂综合。然而,在数学函数领域,我们的研究揭示了一个关键的模态干扰现象:即使是先进的模型,在执行文本计算推理时,也倾向于忽视或误解关键的视觉线索。为应对这一挑战,我们提出了Func-R1,它协同地协调了精确的视觉感知与严谨的逻辑推理。具体而言,基于一种显式解耦的架构,我们采用分层后训练框架,逐步识别关键视觉证据并进行深入的理论推理。此外,我们提出了感知对齐的理论优化(PATO)策略,以引导策略更新朝向内化基本理论属性,同时在推理过程中动态纠正异质视觉信息。跨多个不同基准的大量实验表明,Func-R1在开源MLLMs中实现了最优性能,甚至在MathVerse的面向函数的任务上以8.4%的提升超越了GPT-5。
英文摘要
Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the realm of mathematical functions, our investigation reveals a critical modality interference phenomenon: even advanced models, while performing textual computational reasoning, tend to disregard or misinterpret essential visual cues. To address this challenge, we propose Func-R1, which synergistically harmonizes precise visual perception and rigorous logical reasoning. Concretely, built upon an explicitly decoupled architecture, we employ a hierarchical post-training framework to progressively identify critical visual evidence and conduct in-depth theoretical reasoning. Furthermore, the Perception-Aligned Theoretic Optimization (PATO) strategy is proposed to steer policy updating towards internalizing fundamental theoretical properties while dynamically rectifying heterogeneous visual information throughout the reasoning process. Extensive experiments across diverse benchmarks demonstrate that Func-R1 delivers the optimal performance among open-source MLLMs, even surpassing GPT-5 with an 8.4% improvement on MathVerse's function-oriented tasks.
CommentsAccepted to EMNLP 2026 (2026 Conference on Empirical Methods in Natural Language Processing)