arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29411cs.LGcs.SC

归纳的尽头:整数序列基准中的描述长度难度与记忆差距

Where Induction Runs Out: Description-Length Difficulty and the Memorisation Gap in Integer-Sequence Benchmarks

Sabilashan Ganeshan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究以OEIS整数序列基准为对象,通过MDL参考学习器发现整数序列存在记忆差距,揭示语言模型在该基准中更多依赖记忆而非归纳,并提出MDL可作为无干扰难度信号。

中文摘要 AI 辅助

来自在线整数序列百科全书(OEIS)的整数序列越来越多地被用于基准测试语言模型的数学推理能力。我们使用一个可精确计算的参考学习器——针对P递归(完整)递推类的两部分最小描述长度(MDL),按序列的每个前缀作为项到达时进行评估,来探究这类基准实际测量的内容。由此得出三项发现:第一,MDL难度是参数计数。发现点nd(符号假设击败逐字存储的第一个前缀长度)可被所选算子的阶数和次数的组合可识别边界几乎精确预测,且它与项的大小无关:将斐波那契数列缩放12个数量级,nd保持不变,因为假设需编码自身的初始条件,而大小会抵消。第二,在大规模场景下,我们精心整理的语料库从未产生过的情况出现了:在20000个OEIS序列中,89.98%的在某个前缀上符合递推的序列在全长上不符合任何递推,我们将此称为“荒野”——归纳形成了理论,又失去了它,且从未恢复。第三,按这些MDL层评估三个语言模型时,我们预先注册的假设被推翻:模型不会在MDL报告无理论的地方编造内容,而是进行适当的弃权(不执行);自信的错误发生了反转,集中在容易的层,其中表观能力追踪的是对序列的识别而非对其规则的归纳。因此,源自OEIS的基准主要测量的是记忆,而MDL提供了一种它们目前缺乏的廉价、无干扰的难度信号。代码和数据已公开。

英文摘要

Integer sequences from the On-Line Encyclopedia of Integer Sequences (OEIS) are increasingly used to benchmark mathematical reasoning in language models. We ask what such benchmarks actually measure, using an exactly computable reference learner: two-part minimum description length (MDL) over the class of P-recursive (holonomic) recurrences, evaluated on every prefix of a sequence as terms arrive. Three findings follow. First, MDL difficulty is a parameter count. The discovery point nd, the first prefix length at which a symbolic hypothesis beats verbatim storage, is predicted almost exactly by a combinatorial identifiability bound on the selected operator's order and degree. It is invariant to term magnitude: scaling Fibonacci over twelve orders of magnitude leaves nd unchanged, because a hypothesis must encode its own initial conditions and the magnitude cancels. Second, at scale the learner exhibits a regime our curated corpus could not produce even once: across 20,000 OEIS sequences, 89.98% of those that fit a recurrence on some prefix fit none at full length. We call this the wilderness -- induction acquires a theory, loses it, and never recovers. Third, evaluating three language models on sequences stratified by these MDL regimes refuted our pre-registered hypothesis: models do not confabulate where MDL reports no theory, but hedge appropriately. Confident errors are inverted, concentrating on the easy stratum, where apparent competence tracks recognition of the sequence rather than induction of its rule. OEIS-derived benchmarks therefore substantially measure memorisation, and MDL supplies a cheap, contamination-free difficulty signal they currently lack. Code and data are released.

补充信息

↑