当规范补全出错时:大型语言模型中跃迁的形式化与测量
When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
浏览论文内容
中文总结 AI 辅助
本文针对大型语言模型的“跃迁”进行形式化定义与测量,发现模型在第二步会放弃默认补全完成跃迁,更高难度失败源于推理预算或约束问题,而非回归默认,为相关争论提供了新依据。
中文摘要 AI 辅助
大型语言模型(LLMs)能否完成从证据到新公理系统的溯因跃迁(通常称为“跃迁”),近期引发了大量争论。主流观点认为LLMs在结构上无法完成此类跃迁,而近期研究对其机制和证据均提出了挑战。然而,这场争论仍难以平息,因为该领域至今缺乏对“跃迁”的形式化定义,也缺乏用于检验双方观点的测量方法。在本文中,我们分四步对“跃迁”进行形式化说明,并对第二步进行测量。这四个步骤分别是:确定部分数据的默认补全是什么、何时必须放弃该默认补全、这种放弃是否正确,以及连续跃迁如何复合。具体而言,我们将跃迁实例定义为一个有限扩展问题,附带机器可验证的证明,证明存在正确的补全,该补全在重命名后具有唯一性,且与数据的规范补全不同。规范补全由左右Kan延拓给出,也是模型在无约束条件下生成的结果,因此可作为默认值。我们证明了跃迁实例是适定的,并建立了一个家族定理,无需枚举即可验证任意难度的实例。我们还进一步形式化了跃迁正确的条件以及连续跃迁的复合方式。最后,我们对9个经认证的实例和4个前沿模型进行了测量。在所有248次带约束的试验中,Kan默认率均为0,因此模型在该步骤确实会发生跃迁,每次都会放弃被排除的默认补全。更高难度下的失败源于推理预算耗尽或约束错误,从未源于回归默认补全。这些结果表明第二步并非瓶颈。若存在争议中的(LLMs)无法完成跃迁的情况,那么该缺陷应在于生成约束或构建框架。代码可在以下网址获取:this https URL
英文摘要
Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and empirical evidence. One of the main reasons why the debate remains open is the difficulty of defining the jump precisely enough to test it. In this paper, we attempt to develop a formal account of the jump in four steps and measure the second. These steps ask what the default completion of partial data is, when the constraints exclude it, whether the new structure agrees with later observations, and how successive jumps compound. We define a \emph{jump instance} as a finite extension problem whose constraints exclude the canonical completions given by the Kan extensions and leave one correct completion up to renaming. In this setting, a model with a canonical default performs the second step by producing the correct completion under the constraints. We evaluate fourteen models across three certified families. The canonical completion returns once in $13{,}300$ constrained answers across all runs. Several calibrated models also give the correct completion reliably, including three API models that solve $159$ of $162$ primary chain trials, suggesting that they can jump at this step. We further formalize the third and fourth steps, whose empirical evaluation remains future work. We hope our work paves the path for formalizing and measuring the full jump in the future. The code of the paper is available at https://github.com/EEthanShi/kan-jump-test.
发表机构
- University of Cambridge(剑桥大学)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。