arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20520cs.AIcs.HCcs.PL

数学问题解决中大型语言模型在可执行推理约束下的表示鲁棒性

Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving

Sagnik Nath, Edith Aurora Graf, Liang Zhang, Diego Zapata-Rivera

首次发表
浏览论文内容

中文总结 AI 辅助

研究大型语言模型在数学问题解决中表示鲁棒性,通过改变问题表面表示评估五个模型,发现其对表示敏感,代码增强未统一提升鲁棒性,揭示了正确性等方面新权衡,强调表示应作为接口设计变量。

中文摘要 AI 辅助

大型语言模型(LLMs)越来越多地在数学问题解决上接受评估,但先前工作常将表示等效的公式视为可互换的,并将推理错误与接口故障混为一谈。本文通过系统改变相同潜在问题的表面表示,包括应用题、文字方程、符号方程和同构释义,研究基于LLM的数学问题解决中的表示鲁棒性。使用精心策划的数学等效问题数据集,在直接答案生成条件下评估五个当代LLMs。发现模型在等效公式间正确性频繁变化,同构重新表述下存在系统衰退。评估代码增强条件时发现,虽能揭示一些模型潜在推理能力,但未统一提高鲁棒性,失败在交互层转移,即使可执行推理成功,表示敏感性常持续。结果表明推理支架未消除表示脆性,而是揭示了正确性、可靠性、延迟和成本间新权衡。我们认为在LLM评估和部署中,应将表示视为一级接口设计变量,尤其对于人工智能辅助问题解决系统。

英文摘要

Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failures. This paper investigates representation robustness in LLM-based mathematical problem solving by systematically varying surface representations of the same underlying problems, including story problems, word-equations, symbolic equations, and isomorphic paraphrases. Using a curated dataset of mathematically equivalent problems, we evaluate five contemporary LLMs under a direct answer generation condition. We find substantial representational sensitivity: models frequently change correctness across equivalent formulations, with nontrivial flip rates across story, symbolic, and word-equation variants. We also observe systematic regressions under isomorphic reformulations, showing that even subtle paraphrase-level changes can degrade performance despite preserved mathematical structure. We then evaluate a code-augmented condition in which models externalize reasoning as executable Python code that is run locally for validation. This interface reveals strong latent reasoning capability in some models that perform poorly under direct prompting, but it does not uniformly improve robustness. Instead, failures shift across interaction layers, from opaque reasoning errors to protocol violations and execution failures. Even when executable reasoning succeeds, representation sensitivity often persists. Overall, our results show that reasoning scaffolds do not eliminate representational brittleness, but expose new tradeoffs among correctness, reliability, latency, and cost. We argue that representation should be treated as a first-class interface design variable in LLM evaluation and deployment, especially for AI-assisted problem-solving systems.

补充信息

↑