arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当剪枝遇上可解释性:在大语言模型中保留稀疏自编码器的鲁棒性

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili

arXiv 2608.25941首次发表:更新:

发表机构

The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨剪枝对大语言模型稀疏自编码器的影响,提出扰动能量理论,指出激活感知剪枝方法更鲁棒,发现中间层更敏感,提出分层稀疏策略并经多模型实验验证。

AI 中文摘要

稀疏自编码器(SAEs)被广泛用于解释大语言模型(LLMs)的内部表征,然而其在事后模型压缩下的可靠性却鲜为人知。我们系统性研究了剪枝如何影响SAE行为,并从理论上证明:对于固定的SAE,其影响由扰动能量(一种协方差加权范数)决定。该视角揭示了幅度剪枝的关键局限:因忽略激活几何,它会扭曲学习到的表征空间并降低SAE功能。相比之下,Wanda和SparseGPT等激活感知方法会隐含控制扰动能量,因此在保留SAE行为方面显著更鲁棒。我们进一步发现所有剪枝方法存在一致的结构脆弱性:中间层比早期或后期层对剪枝显著更敏感。基于此洞见,我们提出分层稀疏分配策略,在相同平均剪枝稀疏度下实现更低困惑度。在四种模型架构上的实验验证了我们的理论发现。代码公开于此https URL。

英文摘要

Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm. This perspective exposes a key limitation of magnitude pruning: by ignoring activation geometry, it distorts the learned representation space and degrades SAE functionality. Activation-aware methods such as Wanda and SparseGPT, in contrast, implicitly control perturbation energy and are therefore substantially more robust at preserving SAE behavior. We further reveal a consistent structural vulnerability across all pruning methods: middle layers are significantly more sensitive to pruning than early or late layers. Guided by this insight, we propose a layer-wise sparsity allocation strategy, achieving lower perplexity under the same average pruning sparsity. Experiments across four model architectures validate our theoretical findings. Code is publicly available at https://github.com/osu-srml/sae-robustness-under-pruning/tree/main.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑