RePro:用于可靠评估大语言模型数学问题求解能力的证明验证基准改写框架
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
- Nanyang Technological University(南洋理工大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- INSAIT
- Sofia University “St. Kliment Ohridski”(索非亚大学“圣·克利门特·奥赫里德斯基”)
- Shenzhen Loop Area Institute(深圳河套学院)
- AIRS
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
RePro是首个整合Lean导向神经自动定理证明器的证明验证基准改写框架,在GSM8K和MATH上实现改写实例100%正确,揭示部分LLM数学求解性能存在记忆效应。
AI中文摘要:
数据污染会破坏大语言模型(LLMs)在数学问题求解任务上的可靠评估。尽管基于改写的评估方法能够缓解模型对训练数据的记忆问题,但现有方法无法保证问题的有效性和答案的正确性。我们提出了证明验证基准改写框架RePro,这是首个将Lean导向的神经自动定理证明器(ATPs)整合到基准改写中的框架,该框架通过Lean验证的证明来改写问题并重新生成答案,确保答案的正确性。在GSM8K和MATH数据集上的实验表明,RePro保留的改写实例达到了100%的明确定义性、可行性和答案正确性,而现有方法仍会产生无效或不正确的实例。此外,多个模型在经过证明验证的改写基准上表现出准确率下降,这表明它们的性能对表面和结构变化敏感,可能部分反映了记忆效应。我们的源代码和数据可在指定URL获取。
英文摘要:
Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctness. We propose Proof-Verified Benchmark Rewriting (RePro), the first framework to integrate Lean-oriented neural automated theorem provers (ATPs) into benchmark rewriting, which rewrites problems and regenerates answers with correctness ensured by Lean-verified proofs. Experiments on GSM8K and MATH show that RePro's retained rewritten instances achieve 100% well-definedness, feasibility, and answer correctness, while existing methods still produce invalid or incorrect instances. Moreover, several models exhibit accuracy drops on proof-verified rewritten benchmarks, suggesting that their performance is sensitive to surface-level and structural variations and may partly reflect memorization effects. Our source code and data are available at https://github.com/AI4Engi/RePro.