arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

求解不等于绘图:奥林匹克几何图解推理基准

Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry

Hsien Xin Peng, Anthony Kim, Alvin Li, Calvin Supasanya, Shivank Garg, Kevin Zhu

arXiv 2608.18111首次发表:更新:

AI 中文总结

该研究针对基础模型图解推理能力评估缺口,构建含954道奥林匹克几何题的基准,发现模型求解与绘图能力存在显著差距,绘图平均编译成功率仅36.14%。

AI 中文摘要

GPT、Claude等基础模型如今已能以出色的熟练度求解奥林匹克级数学问题,几何问题求解也因此成为衡量其数学推理能力的标准代理任务。然而,求解几何问题与绘制问题所依赖的图形并非相同的技能:解题进展往往取决于带有正确辅助构造和关联关系的忠实图形,而能够通过推理得出答案的模型是否也能绘制出这样的图形,目前尚不明确。包括MathVista和MathVerse在内的日益增多的基准,仅衡量模型是否得出正确答案,但据我们所知,没有任何基准单独衡量构建图形本身的独特能力,导致该能力未被评估。我们针对这一缺口推出了一个开源基准:包含954道独立的奥林匹克几何问题,其中有297道题构成的难例子集,每道题都配有解答、人类编写的高保真图形(以可渲染的Asymptote代码呈现),以及一套基于文本、代码、图像、视觉语言模型(VLM)和约束条件的指标,用于衡量我们所说的图解推理能力。对当前基础模型的评估显示,求解与绘图之间存在显著差距:它们生成的图形忠实度明显较低,平均编译成功率仅为36.14%。我们发现,强大的数学推理能力并不意味着具备构建准确几何图形的能力。我们的基准和数据集可在该https链接获取。

英文摘要

Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for their mathematical reasoning. Yet solving a geometry problem and drawing the figure it depends on are not the same skill: progress often hinges on a faithful diagram with the right auxiliary constructions and incidences, and it is unclear that a model which reasons its way to the answer can also produce one. A growing collection of benchmarks, including MathVista, and MathVerse, measures whether models reach the correct answer, but to our knowledge, none isolate the distinct ability to construct the diagram itself, leaving this capability unmeasured. We introduce an open-source benchmark that targets this gap: 954 self-contained olympiad geometry problems, with a 297-problem hard subset, each paired with its solution and a human-authored, high-fidelity diagram in renderable Asymptote code, together with a suite of text-, code-, image-, VLM-, and constraint-based metrics for what we term diagrammatic reasoning. Evaluating current foundation models reveals a pronounced gap between solving and drawing: their diagrams are markedly less faithful, with an average compile success rate of only 36.14\%. Strong mathematical reasoning, we find, does not imply the ability to construct accurate geometric diagrams. Our benchmark and dataset can be accessed at https://huggingface.co/datasets/max98765/hard_geometry_problems_with_diagrams.

CommentsICML 2026, AI4MATH Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑