arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Transformer前馈神经网络神经元的稀疏层间依赖关系

Sparse Inter-Layer Dependencies of Transformer FFN Neurons

Johannes Knittel, Hanspeter Pfister

arXiv 2607.11990首次发表:更新:

AI 中文总结

研究Transformer中FFN模块内部结构难解释问题,引入无需训练的归因方法,发现其神经元层面存在稀疏且结构化的层间依赖关系,该方法为电路级可解释性提供工具并识别候选稀疏路径。

AI 中文摘要

前馈网络(FFN)模块在Transformer架构的参数和计算中占很大比例,但其内部结构因残差流引起的加性叠加而难以解释。我们研究FFN神经元的激活是否能用一组稀疏的先前神经元激活和注意力输出解释。我们引入一种无需训练的归因方法来估计上游神经元和注意力输出对目标神经元激活的相对影响。通过实验发现,当用平均值掩盖其余输入时,少量先前激活和注意力输出子集就能高保真保留神经元激活。考虑上游层固有激活稀疏性时,有效稀疏性更大。同时在所有层应用特定神经元掩码,在适度稀疏水平下模型困惑度基本不变。这些结果表明,尽管参数密集,FFN在神经元层面呈现稀疏且结构化的层间依赖关系。我们的方法为电路级可解释性提供实用、可扩展工具,并识别出对高效推理有潜在影响的候选稀疏路径。

英文摘要

Feedforward network (FFN) blocks account for a large fraction of the parameters and computation in Transformer architectures, yet their internal structure remains difficult to interpret due to the additive superposition induced by the residual stream. We examine whether the activation of an FFN neuron can be explained by a sparse set of preceding neuron activations and attention outputs. We introduce a training-free attribution method that estimates the relative influence of upstream neurons and attention outputs on a target neuron's activation. Empirically, across models and layers, we find that small subsets of preceding activations and attention outputs suffice to preserve neuron activations with high fidelity when all remaining inputs are masked with their average values. Effective sparsity is even greater when accounting for the inherent activation sparsity of upstream layers. Moreover, applying the neuron-specific masks in all layers simultaneously, such that the induced deviations propagate through the network, leaves model perplexity largely unchanged at moderate sparsity levels. These results demonstrate that, despite dense parameterization, FFNs exhibit sparse and structured inter-layer dependencies at the neuron level. Our method provides a practical, scalable tool for circuit-level interpretability and identifies candidate sparse pathways with potential implications for efficient inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑