Transformer MLP 门控阈值是与携带参考方向的耦合
Transformer MLP Gate Thresholds Are Couplings to a Carried Reference Direction
AI总结:
本文证明Transformer MLP门控阈值是残差流携带参考方向的耦合,该方向承载功能,偏置贡献小,且机制在多数架构中普遍存在。
AI中文摘要:
Transformer 残差流的语料平均方向是一个在所有输入间共享的组件,在表征分析前通常通过均值中心化移除。我们提出证据表明它是一个功能性组件:MLP 门控群体设定其工作点所依据的参考。在 Phi-2 中,对静息门控预激活的精确分解显示,在堆栈中层,超过 99.9% 的门控的静息抑制由与携带平均方向的耦合 $w \cdot b_L$ 承担,其强度为显式偏置参数的 48-56 倍,且该耦合具有方向特异性:在匹配范数的随机方向上,群体放电率的排序 Spearman $\rho \leq 0.09$,而参考方向达到 0.95。(该 0.95 本身近乎同义反复;论文推导了其零假设。)在因果层面,移除流在该参考上的投影会使超阈值放电增加约 9 倍,呈剂量单调性,为范数匹配对照的 23-59 倍;随机初始化的孪生模型则平坦,替换测试表明方向承载功能而幅度不承载。方向匹配对照也印证了这一点:沿参考注入噪声的代价是沿随机方向相同能量的 8-42 倍。该分解在另外四个涵盖两种门控类型(带显式门控偏置的 GELU、无偏置的 SwiGLU)的家族上复现,剂量反应在其中两个上复现。在八模型扫描中,该机制存在于每个 GELU 和 SiLU 家族中,仅在 OPT 中缺失,其中对立的 LayerNorm 偏置抵消了携带的参考。门控阈值被实现为与网络为其所读取分布构建的常数的耦合,偏置参数贡献甚微;该常数中有多少携带在流中、多少在参数中取决于架构。
英文摘要:
The corpus-mean direction of a transformer's residual stream is a component shared across all inputs, and is commonly removed by mean-centering before representational analysis. We present evidence that it is a functional component: the reference against which the MLP gate population sets its operating point. In Phi-2, an exact decomposition of resting gate pre-activations shows that at mid-stack layers the resting inhibition of >99.9% of gates is carried by the coupling $w \cdot b_L$ to the carried mean direction, at 48-56$\times$ the explicit bias parameter, and the coupling is direction-specific: a random direction at matched norm orders the population's firing rates at Spearman $ρ\leq 0.09$ where the reference reaches 0.95. (That 0.95 is near-tautological on its own; the paper derives its null.) Causally, removing the stream's projection on the reference multiplies above-threshold firing by about 9$\times$, dose-monotonically, at 23-59$\times$ a norm-matched control; a random-initialised twin is flat, and replacement tests show that the direction carries the function and the magnitude does not. A direction-matched control makes the same point: noise injected along the reference costs 8-42$\times$ the same energy along a random direction. The decomposition replicates on four further families spanning both gate types (GELU with an explicit gate bias, bias-free SwiGLU), and the dose-response on two of them. Across an eight-model scan the mechanism is present in every GELU and SiLU family and absent only in OPT, where an opposing LayerNorm bias cancels the carried reference. Gate thresholds are implemented as couplings to a constant the network builds for the distribution it is reading, with the bias parameters contributing little; how much of that constant is carried in the stream and how much in parameters depends on the architecture.