发表机构
University of Moratuwa(莫拉图瓦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控损坏框架揭示,数学智能体在无验证时准确率从100%降至72.4%,强制反思可恢复至100%,验证策略与验证器可用性共同决定可靠性。
AI 中文摘要
数学问题求解通常需要确定性的计算步骤,智能体将这些步骤委托给工具并隐式信任它们。然而,工具可能无声地失败,返回看似合理但不正确的结果。智能体能在多大程度上检测和纠正被篡改的工具调用输出?我们通过一个受控的损坏框架来研究这一问题,其中隐藏的拦截器在针对性的问题上用看似合理的错误信息替换工具调用结果。我们在31个问题上,在四种验证设计下评估了智能体,包括无验证(基线)、强制同上下文反思、可选新上下文验证和可选结构验证。在没有验证的情况下,损坏导致准确率大幅下降,从100%降至72.4%。强制反思完全恢复了这一性能,达到100%。可选验证仅在模型主动调用时提高准确率。我们的结果表明,检查频率与鲁棒性差异密切相关,而不平等的调用阻止了对验证器质量的受控比较。一项支持性恢复实验表明,在明确检测后,完全重启问题在100%的情况下成功。这些发现表明,验证器的可用性和验证策略是数学智能体可靠性的两个独立组成部分。强制策略强制执行验证,而可选策略则依赖于模型自身的选择来调用它。
英文摘要
Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model's own choice to invoke it.
Comments8 pages, 2 figures, 3 tables. Accepted to the 6th Workshop on Mathematical Reasoning and AI (MathAI) at NeurIPS 2026