大型语言模型中的数字能力:基本局限与改进路径
Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
浏览论文内容
中文总结 AI 辅助
本研究针对LLMs在基础数值任务中的不可靠性,提出数值基础框架(NGF)并结合三大基准评估前沿模型,给出改进数值能力的部署建议与研究方向。
中文摘要 AI 辅助
大型语言模型(LLMs)在数学推理基准测试中表现出色,但在基础数值任务上仍不可靠,包括量级比较、大整数算术、分数和科学计数法。本综述将基础数值理解作为一种与高级数学推理不同的能力进行研究,提出了数值基础框架(Numerical Grounding Framework, NGF),该框架将数字能力分解为表征基础(Representational Grounding, RG,即数字形式到数值、量级及等价表征的映射)和过程基础(Procedural Grounding, PG,即按照数学定义执行算术运算)。我们利用NGF梳理了近期诊断基准、失败模式、结构解释及缓解策略,综述了与分词、位置编码、嵌入几何和预训练数据分布相关的证据;还将NGF应用于对三大前沿模型系列的协同评估,涉及Number Cookbook、NumericBench和GSM-Symbolic三个基准,对比了原子、上下文及推理辅助的数字能力。数字感知分词、Abacus Embeddings等架构干预措施可改进从头训练的模型,但预训练系统用户通常无法使用这些方法,对他们而言,监督微调、推理支架和外部工具更实用。最后,我们为基础模型的可靠数值行为提出了部署建议和研究方向。
英文摘要
Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerical understanding as a capability distinct from high-level mathematical reasoning. We propose the Numerical Grounding Framework (NGF), which decomposes numeracy into Representational Grounding (RG), mapping numeral forms to value, magnitude, and equivalent representations, and Procedural Grounding (PG), executing arithmetic operations in accordance with their mathematical definitions. Using NGF, we organize recent diagnostic benchmarks, failure modes, structural explanations, and mitigation strategies. We review evidence concerning tokenization, positional encoding, embedding geometry, and pretraining-data distribution. We also apply NGF in a coordinated evaluation of three frontier model families across Number Cookbook, NumericBench, and GSM-Symbolic, comparing atomic, contextual, and reasoning-assisted numeracy. Architectural interventions such as digit-aware tokenization and Abacus Embeddings can improve models trained from scratch but are generally unavailable to users of pretrained systems, for whom supervised fine-tuning, reasoning scaffolds, and external tools are more practical. We conclude with deployment recommendations and research directions for more reliable numerical behavior in foundation models.
发表机构
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。