arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越答案密钥:用于步骤级数学验证的大语言模型鲁棒性评估

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

Fateme Mazdarani, Carlos Toxtli

arXiv 2608.28725首次发表:更新:

AI 中文总结

本研究构建线性方程基准评估大语言模型的步骤级数学验证鲁棒性,发现开源模型对受扰动等价解题过程的错误拒绝率高,微调等方法可部分提升鲁棒性但存在权衡。

AI 中文摘要

大语言模型(LLMs)越来越多地被用作评分器、验证器和过程审计员,但大多数数学评估仍强调最终答案的准确性,这可能掩盖模型是否能够验证非规范但有效的解题过程。我们引入了一个受控线性方程基准,用于评估作为评估者角色的LLMs,每个实例要求模型判断最终答案正确性、步骤级过程正确性以及第一个错误步骤。我们对最先进的开源LLMs的评估显示存在显著的鲁棒性差距:能准确评估规范解题过程的模型在面对受扰动但逻辑等价的变体时往往表现不佳。在GPT-OSS 20B、Qwen3-14B和Phi-4-Reasoning上,基础模型在规范解题过程上表现良好,但在受扰动解题过程上性能大幅下降,尤其是在错误定位方面。在有效的受扰动解题过程上,基础模型的错误拒绝率达到75.6%-85.3%,显示出对规范解题过程形式的强敏感性。监督微调、蒸馏和测试时计算在某些场景下可提升鲁棒性,但增益依赖于模型,且可能与规范性能形成权衡。结果表明,可靠的过程级验证仍具挑战性,即使在具有精确真值的简单代数领域,评估者鲁棒性也应与求解器准确性分开衡量。

英文摘要

Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can obscure whether a model can verify a non-canonical but valid solution trace. We introduce a controlled linear-equation benchmark for evaluating LLMs in the evaluator role. Each instance asks the model to judge final-answer correctness, step-level trace correctness, and the first incorrect step. Our evaluation of state-of-the-art open LLMs reveals a significant robustness gap: models that accurately evaluate canonical solutions often fail when presented with perturbed but logically equivalent variants. Across GPT-OSS 20B, Qwen3-14B, and Phi-4-Reasoning, base models perform well on canonical traces but degrade substantially on perturbed traces, especially for error localization. On valid perturbed traces, base-model false-rejection rates reach 75.6-85.3%, showing strong sensitivity to canonical solution form. Supervised fine-tuning, distillation, and test-time compute improve robustness in some settings, but gains are model dependent and can trade off against canonical performance. The results show that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.

CommentsAccepted to 2026 IEEE International Conference on Machine Learning and Applications (ICMLA)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑