已定位但不可释放:静默门控反转与有界线性释放
Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release
浏览论文内容
中文总结 AI 辅助
本研究在2570万参数的因果证据判别Transformer模型上,测试了检测、定位与释放潜在结构的流程,发现定位成功但门控反转、线性释放有界,两种失效可分离且未推翻定位结果。
中文摘要 AI 辅助
越来越多的研究表明,语言模型会表示出与任务相关的潜在结构,但却未能加以利用。这种结构一旦被定位,是否能转化为实际行为是一个很少被端到端测试的独立问题。我们将完整的流程——检测、定位与释放——提交给一项完全预先注册的压力测试,该测试在一个2570万参数的Transformer模型上进行,该模型经过因果证据判别训练,此前已有已知的抑制现象(存在潜在因果结构但未在行为中使用)被记录。所有阈值、声明模板和决策树分支在对应数据存在前均已哈希并存档。三项发现如下:(i) 定位成功:对中间层的观测-证据通道进行干预,可在原本被抑制的场景中恢复目标行为(成对释放优势为0.563和0.854,97.5%置信区间不包含0;最佳位点释放率为0.889)。(ii) 门控在分布外失效:一个校准为在0个分布外校准场景触发的检测器,在6.9%-7.3%的保留分布内生成中触发,但在2400个实际需要触发的保留生成中触发数为0——这是一种完全的反转,使门控流程静默地退化为基础模型。(iii) 线性释放有上限:移除门控并无条件注入每个实例的线性方向,会产生单调的剂量反应,但在远低于预先注册的释放阈值(截距0.382→0.311→0.264,阈值≤0.08)时趋于平稳;每个实例的适应性带来的增益小于±0.03。该失效具有双重定位性:检测器在分布外被反转,且该位点和分辨率下的整个线性释放方向族都与充分性存在界限。这两种失效可分离,且均未推翻定位结果。所有数字均可追溯至已发布审计链中的哈希制品。
英文摘要
A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. We submit the complete pipeline -- detect, localize, and release -- to a fully preregistered stress test on a 25.7M transformer trained on causal-evidence discrimination, where a known suppression phenomenon (latent causal structure present but behaviorally unused) has previously been documented. Every threshold, claim template, and decision-tree branch was hashed and archived before any corresponding data existed. Three findings. (i) Localization succeeds: interventions at observation-evidence channels of mid layers restore target behavior on otherwise-suppressed worlds (paired release advantages $0.563$ and $0.854$, 97.5% CIs excluding zero; best-site release rate $0.889$). (ii) Gating fails out of distribution: a detector calibrated to trigger on zero out-of-distribution calibration worlds triggers on 6.9-7.3% of held-out in-distribution generations and on zero of the 2,400 held-out generations that actually need it -- a complete inversion that silently reduces the gated pipeline to its base model. (iii) Linear release is capped: removing the gate and injecting a per-instance linear direction unconditionally yields a monotone dose-response that plateaus far below the preregistered release margin (intercept $0.382 \to 0.311 \to 0.264$ vs. threshold $\le 0.08$); per-instance adaptivity adds less than $\pm 0.03$. The failure is doubly located: the detector is OOD-inverted, and the entire family of linear release directions at this site and resolution is bounded away from sufficiency. The two failures are dissociable, and neither overturns localization. Every number traces to a hashed artifact in the released audit chain.
发表机构
- Tsingjiao Information Science (Beijing) Co., Ltd.(清教信息科学(北京)有限公司)
机构由 AI 辅助整理,请以论文原文为准。