主动SAE特征平面是否具有更多的全同性?Gemma中的预注册反转
Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
浏览论文内容
中文总结 AI 辅助
研究在Gemma 2 2B中全同性是否集中在主动SAE特征平面,通过特定规则测量全同性,预注册相关内容。结果证伪预测,主动特征平面全同性低于混合特征对照,还进行了冻结后诊断,是操作反转非因果声明,原因尚待探究。
中文摘要 AI 辅助
本文测试了在Gemma 2 2B中全同性是否集中在主动稀疏自编码器(SAE)特征平面上,这是更广泛语义集中预测的具体操作化。通过使用仪器的受限雅可比传输规则在小环上携带局部框架,然后通过封闭面积对所得旋转进行归一化,在最终令牌层12到层13的残差流读出中测量全同性。设计、重要性阈值、分析和判定规则在检查分析测量之前进行了预注册和冻结。预测被反向证伪:主动特征平面的全同性低于匹配的混合特征对照,调整后的对数对比度为-0.29439,95%区间为[-0.43989,-0.14889]。在这种设计中,仅幅度解释不成立,而在匹配幅度下,随机、混合特征和主动特征平面之间的三向排序未定义,因为共同支持失败。在相同读出处的冻结后诊断在一个小验证子集上支持面积定律,在简单配对回归下限制匹配中心位移,并将传输失真识别为一种存在的机制或混杂因素。因此,结果是一个狭义的、可审计的操作反转,而不是一个因果声明,即意义抑制全同性。原因仍然未知,激活强度几何、特征参与度、字典几何、匹配中心位移、激活流形接近度和传输剪切是可能的替代原因。
英文摘要
This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the broader semantic-concentration prediction. Holonomy is measured at the final-token layer-12 to layer-13 residual-stream readout by carrying a local frame around small loops using the instrument's restricted-Jacobian transport rule, then normalizing the resulting rotation by enclosed area. The design, materiality threshold, analysis, and verdict rules were preregistered and frozen before the analysed measurements were inspected. The prediction was falsified in reverse: active-feature planes carried less holonomy than matched mixed-feature controls, with an adjusted log contrast of -0.29439 and 95% interval [-0.43989, -0.14889]. A magnitude-only explanation was not supported in this design, while the three-way ordering across random, mixed-feature, and active-feature planes was undefined at matched magnitude because common support failed. Post-freeze diagnostics at the same readout supported the area law on a small validation subset, bounded matched-center displacement under a simple paired regression, and identified transport distortion as a live mechanism or confound. The result is therefore a narrow, auditable operational reversal, not a causal claim that meaning suppresses holonomy. The cause remains open, with activation-strength geometry, degree of feature engagement, dictionary geometry, matched-center displacement, activation-manifold proximity, and transport shear as live alternatives.