AI 中文总结
本研究针对LLM数学答案验证难题,提出AMTFV方法,通过MTF接口解耦验证建模与执行,在5个数学推理数据集上较基线提升最高8.3个百分点,验证复杂度越高增益越显著。
AI 中文摘要
大型语言模型已展现出强大的数学问题求解能力,但可靠验证其候选答案仍具挑战性。现有代表性方法主要通过自然语言反思修正输出,或直接生成验证程序辅助验证;前者可能无法可靠支持精确计算,后者过早地将数学建模与底层实现耦合。我们提出AMTFV(Agentic Mathematical Tool-Flow Verification,智能体数学工具流验证),通过引入数学工具流(Mathematical Tool Flow,MTF)作为中断-执行-恢复接口,将验证建模与具体执行解耦,并通过数学工具箱支持精确计算。具体而言,验证智能体首先构建验证工作流,在MTF请求中编码需可靠执行的数学对象与计算意图,将请求发送至数学工具箱智能体;后者解析请求、生成可执行调用并将其分派至后端进行精确计算。工具输出可用于候选答案裁定、答案修正及验证工作流修正。我们在5个具有挑战性的数学推理数据集上,对来自DeepSeek、GPT和Gemini的7种模型配置评估AMTFV。实验结果表明,AMTFV总体优于本研究中评估的代表性基线;在单个模型配置下,它较最强基线的平均准确率提升最高达8.3个百分点,且在中等和高验证复杂度样本上的增益更大。
英文摘要
Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challenging. Existing representative methods mainly revise outputs through natural-language reflection or assist verification by directly generating verification programs; the former may not reliably support exact computation, whereas the latter prematurely couples mathematical modeling with low-level implementation. We propose AMTFV (Agentic Mathematical Tool-Flow Verification). By introducing Mathematical Tool Flow (MTF) as an interrupt--execute--resume interface, AMTFV decouples verification modeling from concrete execution and supports exact computation through a mathematical toolbox. Specifically, the verification agent first constructs a verification workflow, encodes the mathematical objects and computational intent requiring reliable execution in an MTF request, and sends it to the mathematical toolbox agent. The latter parses the request, generates executable calls, and dispatches them to the backend for exact computation. Tool outputs then support candidate-answer adjudication, answer revision, and verification-workflow revision. We evaluate AMTFV on five challenging mathematical reasoning datasets with seven model configurations from DeepSeek, GPT, and Gemini. Experimental results show that AMTFV outperforms the representative baselines evaluated in this study overall; under an individual model configuration, it improves average accuracy over the strongest baseline by up to 8.3 percentage points, with larger gains on samples of medium and high verification complexity.
Comments19 pages, 9 figures