点火指数:测量语言模型中的全局工作空间动力学
The Ignition Index: Measuring Global Workspace Dynamics in Language Models
浏览论文内容
中文总结 AI 辅助
该研究提出点火指数I作为量化指标,用于测量Transformer语言模型的全局工作空间动力学,揭示了不同架构模型的点火特性及相关规律,为GWT与机制可解释性搭建了定量桥梁。
中文摘要 AI 辅助
我们提出了点火指数(Ignition Index,I),这是一种经过验证的标量度量,用于实现全局工作空间理论(Global Workspace Theory,GWT)的全有或全无点火预测,适用于Transformer语言模型。该度量将四层参数的Sigmoid函数拟合为每一层线性探测准确率随输入信号强度变化的函数,提取陡峭度参数β-hat:高值表示突然的、类似点火的转变,低值表示渐进式积累。在涵盖五个架构族的11个模型中,打乱标签的对照实验显示,其对真实语言结构的选择性比虚假探测能力高9.6倍(p<0.001,Mann-Whitney U检验)。我们发现:(1)前馈Transformer的总β-hat比SSM高89%(p<1e-13,Cohen's d=0.52),其中Mamba表现出近线性曲线,与缺乏全局广播一致;(2)Huginn-3.5B沿其迭代轴的点火量比沿深度轴高2.12倍,表明循环架构在循环维度上表现出类似工作空间的转变;(3)Pythia-410M在训练步骤256处显示出由PELT检测到的相变(+67%),早于诱导头形成;(4)将点火与模型规模和信号强度关联的假设未得到证实,表明Transformer架构可能已饱和可用的点火机制。点火指数首次在GWT的动力学预测与机制可解释性之间建立了经过验证的定量桥梁,具有9.6倍的测量选择性和架构级可区分性,这在之前的缩放文献中未被描述。代码:this https URL
英文摘要
We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all-or-none ignition prediction in transformer language models. The metric fits a four-parameter sigmoid to per-layer linear probe accuracy as a function of input signal strength, extracting steepness parameter beta-hat: high values indicate abrupt, ignition-like transitions; low values indicate graded build-up. Across 11 models spanning five architecture families, shuffled-label controls demonstrate 9.6-fold selectivity for genuine linguistic structure over spurious probe capacity (p < 0.001, Mann-Whitney U-test). We find: (1) Feedforward transformers exceed SSMs by 89% in aggregate beta-hat (p < 1e-13, Cohen's d = 0.52), with Mamba exhibiting near-linear profiles consistent with absent global broadcast. (2) Huginn-3.5B exhibits 2.12-fold higher ignition along its iteration axis than its depth axis, demonstrating that recurrent architectures manifest workspace-like transitions along the recurrence dimension. (3) Pythia-410M shows a PELT-detected phase transition at training step 256 (+67%), preceding induction-head formation. (4) Hypotheses linking ignition to model scale and signal strength were not confirmed, suggesting transformer architectures may saturate available ignition mechanisms. The Ignition Index provides the first validated quantitative bridge between GWT's dynamical predictions and mechanistic interpretability, with 9.6-fold measurement selectivity and architecture-level discriminability not previously characterized in the scaling literature. Code: https://github.com/saman-rahbar/ignition-index
发表机构
- Dialpad, Inc.(戴尔德公司)
机构由 AI 辅助整理,请以论文原文为准。