发表机构
The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨剪枝对大语言模型稀疏自编码器的影响,提出扰动能量理论,指出激活感知剪枝方法更鲁棒,发现中间层更敏感,提出分层稀疏策略并经多模型实验验证。
AI 中文摘要
稀疏自编码器(SAEs)被广泛用于解释大语言模型(LLMs)的内部表征,然而其在事后模型压缩下的可靠性却鲜为人知。我们系统性研究了剪枝如何影响SAE行为,并从理论上证明:对于固定的SAE,其影响由扰动能量(一种协方差加权范数)决定。该视角揭示了幅度剪枝的关键局限:因忽略激活几何,它会扭曲学习到的表征空间并降低SAE功能。相比之下,Wanda和SparseGPT等激活感知方法会隐含控制扰动能量,因此在保留SAE行为方面显著更鲁棒。我们进一步发现所有剪枝方法存在一致的结构脆弱性:中间层比早期或后期层对剪枝显著更敏感。基于此洞见,我们提出分层稀疏分配策略,在相同平均剪枝稀疏度下实现更低困惑度。在四种模型架构上的实验验证了我们的理论发现。代码公开于此https URL。
英文摘要
Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm. This perspective exposes a key limitation of magnitude pruning: by ignoring activation geometry, it distorts the learned representation space and degrades SAE functionality. Activation-aware methods such as Wanda and SparseGPT, in contrast, implicitly control perturbation energy and are therefore substantially more robust at preserving SAE behavior. We further reveal a consistent structural vulnerability across all pruning methods: middle layers are significantly more sensitive to pruning than early or late layers. Guided by this insight, we propose a layer-wise sparsity allocation strategy, achieving lower perplexity under the same average pruning sparsity. Experiments across four model architectures validate our theoretical findings. Code is publicly available at https://github.com/osu-srml/sae-robustness-under-pruning/tree/main.