发表机构
Johannes Kepler University Linz(约翰内斯·开普勒大学林茨分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型推理轨迹的解读偏差,通过重启控制截断探测法和难度控制测试,揭示推理轨迹的结果信息混淆问题,提出需采用问题内评估等方法建立尝试内信息。
AI 中文摘要
大语言模型的推理轨迹被广泛解读为包含“突破”时刻和早期可辨识的结果,这两种解读都依赖于在断言层面缺失反事实控制的测量;我们提供了这两种控制。首先,重启控制截断探测法将解决方案是否符合延续预算与前缀是否具备新鲜计算无法获取的价值区分开来,在匹配的总生成token预算下,比较每个锚点的延续解决率与从头重启曲线。应用于178个问题-模型单元(89个MATH问题×两个小型开放模型,这是一个结果盲但以难度为目标的队列),178个单元中仅有1个作为前缀受限单元留存;重启剂量反应将计算受限模型与能力受限模型区分开来;且无论匹配预算处于重启网格的何处,延续模型自身的前缀都优于重启(9个中的9个)——这主要是计算压缩而非可达性扩展。其次,一项预先注册的、难度控制的测试发现,除了问题难度基线外,早期窗口内部信号中没有可检测的结果信息,对公共语料库的两项无生成分析显示了为何需要这种控制:无轨迹的难度代理在192K DeepSeek-R1生成结果上达到AUROC 0.873(在已发布的探测范围内),对最接近的已发布早期窗口正结果的紧密匹配重建恢复了可比的合并结果(0.849),而在问题内部,其在所有十个锚点处与随机无统计学差异(t=4时为0.496);事后针对性探测仅发现小的平均残差,集中在三个低失败问题上。高合并探测AUROC本身无法建立尝试内信息;需要仅问题基线或问题内评估。
英文摘要
Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the continuation budget from when a prefix carries value that fresh computation cannot buy, comparing per-anchor continuation solve rates against from-scratch restart curves at matched total generated-token budget. Applied to 178 problem-model cells (89 MATH problems x two small open models, an outcome-blind but difficulty-targeted cohort), exactly 1 of 178 cells survives as prefix-limited; restart dose-response separates a compute-starved model from a capability-limited one; and wherever the matched budget lies inside the restart grid, continuing the model's own prefix beats restarting (9 of 9) -- predominantly compute compression rather than expanded reachability. Second, a pre-registered, difficulty-controlled test finds no detectable outcome information in early-window internal signals beyond a problem-difficulty baseline, and two generation-free analyses of public corpora show why this control is needed: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations -- inside the published probe range -- and a closely matched reconstruction of the closest published early-window positive recovers a comparable pooled result (0.849) while within problem it is statistically indistinguishable from chance at all ten anchors (0.496 at t=4); a post-hoc within-targeted probe finds only a small average residual, concentrated in three low-failure problems. High pooled probe AUROCs cannot by themselves establish within-attempt information; a question-only baseline or within-problem evaluation is required.
Comments25 pages, 11 figures, 4 tables. Also available at doi:10.5281/zenodo.22261107. Code and pre-registered protocols: https://github.com/bulutyigit/problem-not-path