arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果结构可诱导但功能解耦:类型化机制库的路由/读出边界

Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library

Xining Xun

arXiv 2608.11767首次发表:更新:

发表机构

Tsingjiao Information Science (Beijing) Co., Ltd.(清教信息科学(北京)有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究在两种规模的语言模型中发现,类型化机制库的因果结构可由监督诱导,其路由功能与答案读出解耦,且无额外开销、可审计恢复,为因果知识组织机制提供了新见解。

AI 中文摘要

当语言模型回答干预性问题时,其所需执行的计算取决于查询所需的证据类型。我们报告了Transformer组织因果知识时的一种解耦现象:由类型级监督诱导的逐槽类型结构负责路由,但与答案读出在功能上解耦。我们通过一个类型化机制库验证了这一点,该库是按证据类型划分的离散机制槽,可在状态级别审计,研究在具有精确干预性真值的因果世界基准上,采用冻结协议,设置两种规模(2260万和1.25亿参数)。我们报告四项预注册发现:(i)起源:逐槽类型结构由类型级监督诱导,在架构相同的无监督对照组中不存在,无法通过无内容的门控标签获得,统计上可归因于监督信号,在1.25亿参数规模下,经强大的预注册协议重复验证,所有9个单元均通过;(ii)边界:诱导的结构是具有明确路由/读出边界的类型化路由索引,槽代码为路由提供支撑,但不驱动答案读出(|Δŷ|≤3.4×10⁻⁶,零附带影响,3个随机种子,在5.6倍规模窗口内稳定),因此我们不提出行为可编辑性主张;(iii)成本:该结构无额外开销,语言模型质量与参数匹配的整体模型相差在0.0082纳特以内;(iv)可信性:库状态在编辑下完全局部,且可逐位精确恢复,每个种子可进行250次单次编辑和1000次堆叠恢复,零失败。我们还发现,无监督对照组本身会随规模变化,因此在某一规模校准的对照组用于另一规模比较时可能存在混淆。所有主张均与预注册的、可机器核查的标准绑定,该标准在其管辖的数据之前已归档,完整审计轨迹(包括一项未通过的标准及冻结协议的处理方式)作为附录发布。

英文摘要

When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer organizes causal knowledge: slot-by-type structure induced by type-level supervision organizes routing, yet remains functionally decoupled from answer readout. We establish this with a typed mechanism library -- discrete mechanism slots partitioned by evidence type, auditable at the state level -- on a causal-world benchmark with exact interventional ground truth, under a frozen protocol, at two scales (22.6M and 125M). Four preregistered findings. (i) Origin. Slot-by-type organization is induced by type-level supervision: absent in architecturally identical unsupervised controls, not buyable by content-free gating labels, and statistically attributable to the supervision signal, replicating at 125M under a powered preregistered protocol (all nine cells passed). (ii) Boundary. The induced structure is a typed routing index with a sharp routing/readout boundary: slot codes scaffold routing but do not drive answer readout ($|Δ\hat{y}| \le 3.4\times10^{-6}$, zero collateral, three seeds, stable across a 5.6x scale window) -- we therefore make no behavioral-editability claim. (iii) Cost. The structure is free: LM quality matches a parameter-matched monolith within 0.0082 nats. (iv) Trust. The library state is exactly local under edit and bit-exactly revertible -- 250 single-edit and 1,000 stacked reverts per seed, zero failures. We further find that the unsupervised null itself moves with scale, so comparisons reusing a null calibrated at one scale may be confounded at another. Every claim is tied to a preregistered, machine-checkable criterion archived before the data it governs; the full audit trail, including one criterion we failed and how the frozen protocol handled it, is released as an appendix.

Comments17 pages, 9 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑