从符号感知到逻辑演绎:引导语言模型进行几何推理的框架
From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该框架通过几何视觉解析器和符号求解器,使纯语言模型在平面几何推理上达到与大型多模态模型相当的性能,并基于2025年中考基准验证了其有效性与可解释性。
AI中文摘要:
平面几何仍然是人工智能领域的一项重大挑战,需要整合视觉感知和数学推理。虽然大型多模态模型(LMMs)天然能处理视觉-语言输入,但它们通常计算密集且不透明。我们证明,一个纯粹的大型语言模型(LLM),在配备专门模块后,可以在复杂的几何问题上与最先进的LMMs相媲美。我们的框架集成了一个几何视觉解析器,它将图表转换为符号形式,以及一个符号求解器,它执行形式化演绎,从而减轻幻觉并促进可解释的推理。为了进行严格的评估,我们从2025年中国中考中精选了一个具有挑战性的问题基准,确保数据的新颖性并测试更深的演绎技能。实验表明,我们的方法在性能上与Gemini 2.5 Pro相当,同时提供更清晰、更类似人类的解决方案。
英文摘要:
Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when equipped with specialized modules, can rival state-of-the-art LMMs on complex geometry problems. Our framework integrates a Geometric Vision Parser, which translates diagrams into symbolic form, with a Symbolic Solver that performs formal deductions, thereby mitigating hallucinations and promoting interpretable reasoning. To enable rigorous evaluation, we curate a benchmark of challenging problems from the 2025 Chinese Zhongkao examinations, ensuring data novelty and testing deeper deductive skills. Experiments demonstrate that our approach achieves performance comparable to Gemini 2.5 Pro while delivering clearer, human-like solutions.