利用分块奇异值分解寻找可用的权重机制
Finding Usable Weight Mechanisms with Tiled SVD
浏览论文内容
中文总结 AI 辅助
该研究提出通过列分块奇异值分解直接从线性位点提取机制挂载的方法,在Gemma-2-2B模型与WikiText-2数据集上验证其有效性,相关代码等资源已公开。
中文摘要 AI 辅助
机制可解释性的主流方法是训练代理字典,如稀疏自编码器,并从最大激活文本中标记特征。这类最佳图谱能识别概念,但该概念存在于学习到的字典中,而非网络权重本身。我们提出通过列分块奇异值分解(SVD)直接从线性位点提取机制挂载:每个挂载是三元组(v, u, σ),分别表示触发、写入和强度,其身份即权重规则。我们用预先注册的套件评估挂载,该套件依据全写入能量提升而非分块局部提升进行判断。在Gemma-2-2B模型与WikiText-2(16384-token子样本)上,对全部7个线性映射评分:残差写入(https://github.com/google/gemma_pytorch, attn.o)在子层后RMS归一化后接受 steer,通过52/52个位点层;其他映射仅接受A/B评分(https://github.com/google/gemma_pytorch 26/26个)。总计:182/182个GO。我们发布了库代码、语料库构建器、实验入口点及单元测试。
英文摘要
The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-activating text. The best such atlases identify con- cepts, but that identity lives in the learned dictionary rather than in the network weights them- selves. We propose extracting mechanism mounts directly from linear sites by column-tiled SVD: each mount is a triple (v,u,σ) read as trigger, write, and strength. Identity is the weight rule. We evaluate mounts with a pre-registered suite judged on full-write energy lift rather than tile-local lift. On Gemma-2-2B with WikiText-2 (16,384-token subsample), all seven linear maps are scored: residual writes (mlp.down, attn.o) receive full A/B/C with steer after post-sublayer RMSNorm and pass 52/52 site-layers; other maps receive A/B only (mlp.gate/attn.q/attn.k/effective mlp.up/attn.v 26/26 each). Aggregate: 182/182 GO. We release library code, the corpus builder, the experiment entrypoint, and unit tests.
发表机构
- Aquin Labs(阿昆实验室)
机构由 AI 辅助整理,请以论文原文为准。