发表机构
Oxford Immune Algorithmics; Oxford University Innovation; London Institute for Healthcare Engineering; King’s College London(牛津免疫算法学公司; 牛津大学创新公司; 伦敦医疗工程研究所; 伦敦国王学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出F-ICL基准,以精确贝叶斯最优标准测量语言模型的上下文内算法推理能力,发现多数模型分布与最优标准存在差距,该差距与准确率无关且不受模型规模缩小,基准已开放。
AI 中文摘要
大型语言模型是进行真正的算法推理还是仅完成模式,这一点难以测试,因为大多数基准缺乏正确归纳推理的 ground truth。我们引入F-ICL,这是一个提供精确 ground truth 的上下文学习基准。使用图灵完备机器F,将其补全对称化为sF以消除输出极性偏差,我们穷尽枚举所有长度$L\le13$的15亿个程序,并在有界通用(Levin–Solomonoff)先验下以闭式形式计算贝叶斯最优后验;模型的评分依据是其输出分布在匹配证据下与最优分布的接近程度。每个任务与其按位补全任务配对,最优分布在这两个任务上的评分相同,因此原始任务与补全任务之间的差距可用于隔离模型的归纳偏差。在涵盖37个开放模型(参数规模0.8B至675B)和四个实验室前沿系统的105种输出配置中,模型回答了高达92%的查询,但46个模型中有45个的分布比按键参考更远离最优分布,其行为仅由可见证据拟合的低阶前缀统计量限定。该参考本身是一种算法混合,由无循环的仅打印机器诱导,因此面板隐含的测度更接近无循环混合而非带循环的最优分布,且与参考机器无关。更新也是非单调的,这是现有理论无法解释的:在这种可实现、无噪声的环境中,贝叶斯理性的已解决集合只能增长,但添加示例会产生6545次已解决到未解决的转变,而获得的转变为13702次。该差距无法用准确率预测(斯皮尔曼相关系数$\rho=-0.19$,$p=0.21$),在暴露分布的输出模式中,差距不会随模型规模或前沿代际缩小,且会因训练后的指令和推理而扩大。F-ICL作为开放、可复现的基准和工具包发布。
英文摘要
Whether large language models perform algorithmic inference or pattern completion is hard to test, because most benchmarks supply answers but no distributional reference for what the shown evidence licenses. F-ICL supplies one exactly: we exhaustively enumerate the 86 million valid programs of length at most 13 on a Turing-complete machine F, complement-symmetrised to remove output-polarity bias, and compute the exact posterior under a declared bounded Levin--Solomonoff prior. It is Bayes-optimal for that stated prior rather than universal, and models are never told it exists, so the score reads the inductive prior their served distribution already encodes. Across 105 serving configurations spanning open models from 0.8B to 675B and frontier systems, models answer up to 92% of queries correctly, yet 45 of the 46 exposing distributions sit farther from the F reference than a keystroke reference. This is not an artefact of task selection: on the bit coordinate, the half the length quota cannot distort, 69 of 80 runs stay below the anchor. Fidelity is inert to scale, which accuracy tracks; continuation improves late without converging; and models un-solve a solved task once per two gains, where the F reference does so once per nine and always repairs it. Because absolute distances are reference-dependent, we prove sequential bounds holding for rival priors: any predictor whose prior gives the reference positive weight has bounded cumulative excess loss, and, in a loss never invoking the reference, any Bayesian mixture giving the realised truth positive mass has a bounded truth-loss budget. On 23,998 trajectories, 86.7% already spend over 10 bits of it. Sequences ending by position nine cannot exclude an arbitrarily large finite constant, so these are lower bounds on what a rival prior must already pay. F-ICL is an open benchmark and toolkit.