arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16118cs.AImath.HO

评估大语言模型的数学能力需要理解数学创造力的多种机制

Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity

Silvère Gangloff

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出数学创造力包含多种不可替代的机制,当前LLMs仅具备部分模式的能力,建议围绕该机制分类而非综合基准评估其数学能力。

中文摘要 AI 辅助

应如何评估大语言模型(LLMs)是否能进行数学发明?本文认为该问题目前界定不足:数学创造力并非单一能力,而是由多种机制上不同的意义建构模式构成,包括对数学实践的反思性内省、从科学中进行类比引入、问题驱动的建构,以及跨遥远领域的联结;此外还有一个贯穿各模式的区分:因观察到模式而追求的意义,与因战略需要而追求的意义,本文通过猜想形成的案例来展开这一区分。这些机制可能不可替代,因此某一机制的能力无法迁移到其他机制。本文通过历史案例研究和当前基于Transformer的系统的架构层面分析,指出当前模型的能力集中在由现有构建块的重组和搜索所塑造的模式中;若该描述成立,其余模式原则上无法实现,而非仅速度较慢——不过该描述是否成立本身是开放的经验性问题。由于AI生成证明的能力提升使证明成本降低(该领域的权威人士现已指出这一转变),数学价值正转向当前系统尚未能实现的模式,对AI数学能力的评估应围绕这一分类组织,而非围绕将其混为一谈的综合基准。

英文摘要

How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecified: mathematical creativity is not one capacity but several mechanistically distinct modes of meaning-making - reflexive introspection on mathematical practice, analogical import from the sciences, problem-driven construction, and the bridging of distant domains - together with a further, cross-cutting distinction between meaning pursued because a pattern was observed and meaning pursued because it is strategically wanted, a distinction I develop through the case of conjecture-formation. These mechanisms are likely non-substitutable, so that competence in one does not transfer to the others. Grounding each in a historical case study and in an architecture-level account of current transformer-based systems, I suggest that today's models concentrate their competence in modes shaped by recombination and search over existing building blocks; if that description holds, the remaining modes are out of reach in principle, not just slower - though whether it holds is itself the open, empirical part. Because proof is getting cheaper as AI improves at generating it - a shift the field's own leading voices are now diagnosing - mathematical value is migrating toward the modes current systems cannot yet perform, and evaluations of AI mathematical ability should be organized around this taxonomy rather than around aggregate benchmarks that conflate it.

发表机构

  • University of Ostrava(俄斯特拉发大学)
  • IRAFM

机构由 AI 辅助整理,请以论文原文为准。

↑