arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26177cs.CLcs.AIcs.LG

Magnitude Profile Pruning:用于Transformer压缩的无校准结构化注意力头剪枝

Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression

Kasun Dewage, Marianna Pensky, Heranga K. Rathnasekara, Suranadi De Silva

首次发表
浏览论文内容

中文总结 AI 辅助

提出无校准的注意力头重要性评分方法MP及GQA变体MP-G,通过权重范数统计异常值检测剪枝,在多个模型上匹配或超越数据依赖方法,实现零成本结构化剪枝。

中文摘要 AI 辅助

注意力头的结构化剪枝为压缩Transformer语言模型提供了一种硬件友好的方式。然而,现有的衡量头重要性的方法需要校准数据、梯度计算或Hessian估计。这些需求增加了额外开销,并使方法依赖于数据。我们的工作提出了Magnitude Profile(MP)评分,这是一种无需训练的头部重要性标准,通过对权重行范数进行统计异常值检测来识别可舍弃的头。投影权重落入总体主体范围内的头被剪除,而表现出异常范数(承载着不成比例的表示能力)的头则被保留。我们的工作进一步提出了MP-G,这是一种处理分组查询注意力(GQA)的变体,它将共享的键值组得分分配给相关的查询头。在五个模型上,以12.5%-50%的头稀疏度在WikiText-2困惑度上进行评估,MP-G在OPT-6.7B上所有稀疏度水平下均取得最佳困惑度(12.5%时为18.46,25%时为27.87,50%时为152.0)。MP-G在RoBERTa-large上,在12.5%和25%稀疏度下也取得了最佳结果,困惑度值分别为7.27和10.28,优于依赖校准的基线方法,包括Wanda-Head、SparseGPT-Head和Gradient-Head。它需要零次前向传播、零校准样本或零梯度计算。在50%稀疏度下,头部剪枝可实现高达16%的参数减少,并节省50%的注意力FLOPs。我们的结果表明,仅基于权重的统计评分可以匹配或超越数据依赖的方法用于结构化头部剪枝,为Transformer压缩提供了一种实用且零成本的标准。

英文摘要

Structured pruning of attention heads provides a hardware-friendly way to compress Transformer language models. However, existing methods for measuring head-level importance require calibration data, gradient computation, or Hessian estimation. These requirements add extra overhead and make the methods depend on the data. Our work presents Magnitude Profile (MP) scoring, a training-free criterion for head importance that identifies dispensable heads through statistical outlier detection on weight row norms. Heads whose projection weights fall within the population bulk are pruned, while heads exhibiting outlier norms, which carry disproportionate representational capacity, are preserved. Our work further gives MP-G, a variant that handles Grouped Query Attention (GQA) by distributing shared key-value group scores across associated query heads. Across five models evaluated on WikiText-2 perplexity at 12.5%-50% head sparsity, MP-G achieves the best perplexity on OPT-6.7B at all sparsity levels (18.46 at 12.5%, 27.87 at 25%, 152.0 at 50%). MP-G also gives the best results on RoBERTa-large at 12.5% and 25% sparsity, with perplexity values of 7.27 and 10.28, outperforming calibration-dependent baselines including Wanda-Head, SparseGPT-Head, and Gradient-Head. It requires zero forward passes, calibration samples, or gradient computation. At 50% sparsity, head pruning yields up to 16% parameter reduction with 50% attention FLOP savings. Our results show that weight-only statistical scoring can match or outperform data-dependent methods for structured head pruning, providing a practical, zero-cost criterion for Transformer compression.

补充信息

↑