SpatialMath: 基于空间认知的符号推理框架用于数学问题解决
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
- Indian Institute of Technology Delhi(印度理工学院德里)
- Indian Institute of Technology Abu Dhabi(印度理工学院阿布扎克)
- Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SpatialMath通过整合空间表示到结构化符号推理链中,提升多模态模型在视觉密集型数学问题中的推理能力。
AI中文摘要:
多模态小到中型语言模型(MSLMs)在整合视觉和文本信息方面表现出强大的能力,但仍面临在视觉理解和数学推理方面的显著限制,尤其是在具有不同视觉融合水平的几何问题中。当前模型在准确分解复杂的视觉输入和将感知连接到结构化推理之间存在困难,导致性能不佳。为了解决这些挑战,我们提出了SpatialMath,一种新的基于空间认知的符号推理框架,旨在将空间表示整合到结构化符号推理链中。SpatialMath采用专门的感知模块从视觉图表中提取空间化表示,捕捉关键的几何结构和空间关系。这些表示随后被系统地融入符号推理链中,促进具有视觉理解的结构化推理。为此,我们引入了MATHVERSE-PLUS,一个包含结构化视觉解释和逐步推理路径的新型数据集,用于视觉密集型数学问题。SpatialMath在视觉密集型设置中显著优于强大的多模态基线模型,其在数据增强的监督微调中实现了高达10个百分点的提升。鲁棒性分析表明,增强的空间表示直接提高了推理准确性,强化了在MSLMs中需要结构化感知到推理管道的必要性。
英文摘要:
Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning, particularly in geometric problems with diverse levels of visual infusion. Current models struggle to accurately decompose intricate visual inputs and connect perception with structured reasoning, leading to suboptimal performance. To address these challenges, we propose SpatialMath, a novel Spatial Comprehension-Infused Symbolic Reasoning Framework designed to integrate spatial representations into structured symbolic reasoning chains. SpatialMath employs a specialized perception module to extract spatially-grounded representations from visual diagrams, capturing critical geometric structures and spatial relationships. These representations are then methodically infused into symbolic reasoning chains, facilitating visual comprehension-aware structured reasoning. To this end, we introduce MATHVERSE-PLUS, a novel dataset containing structured visual interpretations and step-by-step reasoning paths for vision-intensive mathematical problems. SpatialMath significantly outperforms strong multimodal baselines, achieving up to 10 percentage points improvement over supervised fine-tuning with data augmentation in vision-intensive settings. Robustness analysis reveals that enhanced spatial representations directly improve reasoning accuracy, reinforcing the need for structured perception-to-reasoning pipelines in MSLMs.