arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同数量,不同答案:语言模型中的数值表示不变性

Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

Ephraim Atta-Duncan

arXiv 2609.25009首次发表:更新:

AI 中文总结

本研究通过生成3,600个问题和8,600个提示,评估五个语言模型在数值表示变换下的答案不变性,发现规范准确率高但轨道不变性下降,并指出科学计数法解析缺陷和单位转换错误等病理,强调评估接口与推理失败的区别。

AI 中文摘要

数值等价的应用题,无论数量以小数、分数、百分比、数字单词、科学计数法还是精确转换的单位表示,都应产生相同的规范答案。我们生成了3,600个精确有理数问题和8,600个提示,涵盖五个保持恒等性的变换族,并评估了五个开放权重系统。经过固定语法审计(无需LLM法官即可规范化常见答案形式)后,规范准确率为0.969-0.996,但轨道正确率降至0.848-0.981,轨道不变性降至0.851-0.981;不变但错误的轨道最多占0.003。大部分严格的解析器崩溃源于乘法形式的科学计数法不在所实现的数字语法中,这说明了评估器接口如何可能伪装成推理失败。一个独特的语义病理仍然存在:Mistral Small 4在单位转换输入上得分为0.699,并产生265个与标签相差精确十的幂次的错误。在另一个独立的9,000次调用实验中,向比较臂分配相等的调用,表示共识在低错误子集上并未优于释义共识,并产生了显著更多的误报。随附的辅助档案包含冻结的基准、评估和审计记录、共识原始响应、清单、分析代码以及一键式论文构建。

英文摘要

Numerically equivalent word problems should yield the same canonical answer whether a quantity is written as a decimal, fraction, percentage, number word, scientific notation, or an exactly converted unit. We generate 3,600 exact-rational problems and 8,600 prompts spanning five identity-preserving transformation families, and evaluate five open-weight systems. After a fixed syntax audit that normalizes common answer forms without an LLM judge, canonical accuracy is 0.969-0.996, but orbit correctness falls to 0.848-0.981 and orbit invariance to 0.851-0.981; invariant-but-wrong orbits account for at most 0.003. Most of the broad strict-parser collapse arises because multiplication-form scientific notation lies outside the implemented number grammar, illustrating how evaluator interfaces can masquerade as reasoning failures. A distinct semantic pathology remains: Mistral Small 4 scores 0.699 on unit-converted inputs and produces 265 errors differing from the label by exact powers of ten. In a separate 9,000-call experiment that allocates equal calls to the compared arms, representation consensus does not outperform paraphrase consensus on a low-error subset and produces substantially more false alarms. The accompanying ancillary archive contains the frozen benchmark, evaluation and audit records, consensus raw responses, manifests, analysis code, and a one-command paper build.

Comments15 pages, 2 figures. Ancillary archive includes the frozen benchmark, evaluation and audit records, consensus raw responses, manifests, analysis code, and tests

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑