权重读取与写入特征:基于激活空间的可扩展参数分解
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
- University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出激活支持参数分解(ASPD),联合分解激活与参数空间,利用内部激活约束权重分解,实现大语言模型可扩展、可解释且因果可编辑的参数分析,并在Qwen-3-8B上验证了机制恢复能力。
AI中文摘要:
激活空间和参数空间为模型计算提供了互补的视角。激活表示信息,而权重读取、变换并写入这些信息。然而,现有的可解释性方法大多分别研究这两个空间,导致表示信息与参数级计算之间的联系未被充分探索。我们引入了激活支持参数分解(ASPD),该方法联合分解激活空间和参数空间,并将每个学习到的权重组件基于其读取或写入的激活特征进行落地。这种落地利用模型内部激活约束了原本非唯一的参数分解,同时内部重建目标为被分析的权重矩阵提供了局部学习信号。这些特性共同使得在预训练大语言模型中实现可扩展、可解释且因果可编辑的参数分解成为可能,并在Qwen-3-8B上进行了演示。学习到的读写组件还可以组合成参数级机制回路。我们使用ASPD来恢复经典IOI回路背后的机制,并追踪通过模型权重的语义变换。
英文摘要:
Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes. This grounding constrains otherwise non-unique parameter decompositions using the model's internal activations, while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed. Together, these properties enable scalable, interpretable, and causally editable parameter decomposition in pretrained large language models, demonstrated on Qwen-3-8B. The learned read--write components can also be composed into parameter-level mechanism circuits. We use ASPD to recover mechanisms underlying the classic IOI circuit and trace semantic transformations through model weights.