发表机构
Tsingjiao Information Science (Beijing) Co., Ltd.(清交信息科学(北京)有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究以Pythia套件为对象,发现内部表征先具备可读性后才产生因果性,提出滞后耦合结构,警示勿从探测器精度推断可操控性,确立表征形成快于因果读出的发展瓶颈。
AI 中文摘要
在整个Pythia套件(160M-12B,8个检查点,4个任务族)中,线性探测器在所有规模下最早可在1000步时从残差流中读取目标变量,但沿同一读取方向进行操控在48个模型-检查点单元中有43个等效于弃权(不执行)。内部可读性系统性地超过因果效能,且滞后性不会随规模缩小。我们将这种结构分解为三条可分离的轨迹:(i)内部可读性,从首个检查点起在所有位置饱和(AUROC≥0.990);(ii)行为可读性,其发展随规模增大而逐渐滞后(12B仅在最终检查点达到0.909);(iii)因果效能,几乎总是等效于弃权(不执行),早期偶尔产生反效果,存在一个孤立的正脉冲(12B,8000步,z=+2.49),我们的网格无法解析。排序以“先读后写”为主(11个单元中11个,无反转)。沿探测器方向的表征余量随训练和规模增长至57倍,而因果写入始终低于余量的0.11%——该变量被不断写入表征却被读出机制忽略。在完全预注册的协议下,两个单 onset 假设均判定为不确定(规模斜率+0.24,95%置信区间[-0.60, +0.87];时间投票3:3),这是三条轨迹分解得出的严谨负结果。预注册的OLMo-2复现保留了该方向但幅度减弱。我们的结果警示勿从探测器精度推断可操控性,并确立了一个发展瓶颈:表征形成可靠地快于因果读出巩固。
英文摘要
Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as early as step 1,000 at every scale -- yet steering along that same reading direction remains null-equivalent in 43 of 48 model-checkpoint cells. Internal readability systematically outruns causal efficacy, and the lag does not shrink with scale. We call this structure lagged coupling and decompose it into three dissociable tracks: (i) internal readability, saturated (AUROC >= 0.990) from the first checkpoint everywhere; (ii) behavioral readability, which develops gradually and progressively later at larger scales (12B reaches 0.909 only at the final checkpoint); (iii) causal efficacy, almost always null-equivalent, occasionally counterproductive early, with one isolated positive pulse (12B, step 8,000, z = +2.49) our grid cannot resolve. The ordering is dominantly read-before-write (11/11 units, no inversion). Representation headroom along the probe direction grows up to 57x with training and scale while causal write-in stays below 0.11% of headroom -- the variable is increasingly written into the representation and increasingly ignored by the readout. Under a fully pre-registered protocol, both single-onset hypotheses resolve INDETERMINATE (scale slope +0.24, 95% CI [-0.60, +0.87]; time vote 3:3) -- a disciplined negative explained by the three-track decomposition. A pre-registered OLMo-2 replication preserves the direction at attenuated magnitude. Our results caution against inferring steerability from probe accuracy and establish a developmental bottleneck: representation formation reliably outpaces causal readout consolidation.
Comments15 pages, 5 figures, 7 tables. Pre-registered developmental interpretability study on the full Pythia suite with an OLMo-2 replication