用于高效电路提取的稀疏权重分解
Sparse Weight Decomposition for Efficient Circuit Extraction
浏览论文内容
中文总结 AI 辅助
提出Sparse Weight Decomposition(SWD)方法,将预训练Transformer权重矩阵分解为稀疏因子以提供可解释电路单元,在低数据量下实现高保真电路提取,适用于多模型及全模型替换,还支持零数据变体以拓展机械可解释性分析应用。
中文摘要 AI 辅助
密集型预训练Transformer无法自然提供可解释单元用于电路提取。现有方法通过学习辅助稀疏表示或训练稀疏模型获取此类单元,会产生大量额外计算,且可能在被分析的表示与原始预训练模型之间引入保真度差距。我们提出Sparse Weight Decomposition(SWD,稀疏权重分解),该方法通过将每个权重矩阵分解为两个稀疏因子,对预训练线性投影进行参数化,其共享的中间坐标可作为可单独寻址的电路单元。无需训练单独的替换网络,此参数化表示支持与学习稀疏特征的方法相同的评分、选择和消融电路提取工作流程。在单矩阵替换中,SWD达到了Transcoder及其他强基线方法实现的保留保真度,同时训练替换模型所用数据量不到这些基线方法的1%。在匹配的替换保真度下,SWD在GPT-2、Qwen2.5和Qwen3.5-27B的任务中,以更少的活跃读写边和所选单元达到相同的电路充分性和必要性目标。我们进一步表明,在对非零因子值进行微调后,SWD对所有注意力和MLP权重矩阵的全模型替换仍然有效。最后,SWD还具有零数据变体,可促进机械可解释性分析(如每步分析)的更广泛应用。
英文摘要
Dense pretrained transformers do not naturally expose interpretable units for circuit extraction. Existing approaches obtain such units by learning auxiliary sparse representations or training sparse models, incurring substantial additional computation while potentially introducing a fidelity gap between the representation being analyzed and the original pretrained model. We propose Sparse Weight Decomposition (SWD), which reparameterizes pretrained linear projections by factorizing each weight matrix into two sparse factors whose shared intermediate coordinates serve as individually addressable circuit units. Without training a separate replacement network, this parametric representation supports the same scoring, selection, and ablation circuit extraction workflow used for methods that learn sparse features. Across single-matrix replacements, SWD matches the held-out fidelity achieved by Transcoder and other strong baselines while using less than 1% of the data that those baselines use to train their replacements. For matched replacement fidelity, SWD reaches the same circuit sufficiency and necessity targets with fewer active read/write edges and selected units across tasks on GPT-2, Qwen2.5, and Qwen3.5-27B. We further show that SWD remains effective for full-model replacement of all attention and MLP weight matrices after fine-tuning the nonzero factor values. Finally, SWD also features a zero-data variant, allowing broader use of mechanistic interpretability analysis (e.g., per-step analysis).