arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NL2AGBench:面向AlphaGeometry的大语言模型自动形式化基准测试

NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

Samuel Xiao, Judy Song, Rory Hu, Ziliang Zong

arXiv 2608.28481首次发表:更新:

发表机构

Valley Christian High School; Vandegrift High School; Groton School; Texas State University(谷基督高中; 范德格里夫特高中; 格罗顿学校; 德克萨斯州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出NL2AGBench基准,评估LLMs将英文几何问题转换为AlphaGeometry兼容形式化表示的能力,发现闭源模型可实现80%以上可执行翻译率,开源模型表现欠佳,还提出错误分类法并验证了多种缓解策略的有效性。

AI 中文摘要

近期,大型语言模型(LLMs)在自然语言理解与数学推理方面展现出强大能力,但其将非形式化数学问题转换为形式化表示的能力仍未得到充分探索。这一局限对于神经符号几何系统(如AlphaGeometry)尤为关键,因为其定理证明引擎需要专用领域特定语言(DSL)的输入。尽管AlphaGeometry达到了接近IMO金牌得主的性能,但将自然语言问题手动转换为其形式化语法仍是显著的可用性瓶颈。为应对这一挑战,我们引入自然语言到AlphaGeometry基准测试(NL2AGBench),用于评估LLMs将英文几何问题转换为AlphaGeometry兼容形式化表示的能力。NL2AGBench在AlphaGeometry内部采用基于执行的验证来评估翻译质量,而非仅依赖文本相似度。我们评估了十个不同参数规模的最先进开源与闭源LLMs,分析可执行翻译准确率、语法正确性及错误特征。实验揭示闭源与开源模型间存在显著性能差距:领先的闭源模型可实现80%以上的可执行翻译率,而即使是最大规模的开源模型也难以持续保留几何约束并生成有效形式化表示。我们引入区分语法与逻辑错误的错误分类法,并研究包括少样本提示、微调及人工引导提示在内的缓解策略,这些策略在多个模型家族中均产生了可测量的改进。

英文摘要

Recent advances in large language models (LLMs) have demonstrated strong capabilities in natural language understanding and mathematical reasoning. However, their ability to translate informal mathematical problems into formal representations remains underexplored. This limitation is particularly important for neuro-symbolic geometry systems such as AlphaGeometry, whose theorem-proving engine requires inputs in a specialized domain-specific language (DSL). Although AlphaGeometry achieves near-IMO gold-medalist performance, manually converting natural-language problems into its formal syntax remains a significant usability bottleneck. To address this challenge, we introduce the Natural Language to AlphaGeometry Benchmark (NL2AGBench), which evaluates LLMs in translating English geometry problems into AlphaGeometry-compatible formal representations. NL2AGBench uses execution-based verification within AlphaGeometry to assess translation quality rather than relying solely on textual similarity. We evaluate ten state-of-the-art open- and closed-source LLMs across multiple parameter scales and analyze executable translation accuracy, syntactic correctness, and error characteristics. Our experiments reveal a substantial performance gap between closed- and open-source models: leading closed-source models achieve executable translation rates above 80%, while even the largest open-source models struggle to consistently preserve geometric constraints and produce valid formalizations. We introduce an error taxonomy distinguishing syntax and logic errors and investigate mitigation strategies, including few-shot prompting, fine-tuning, and human-guided hinting, which yield measurable improvements across multiple model families.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑