发表机构
Florida State University; University of Connecticut; College of William & Mary; Barcelona Supercomputing Center(佛罗里达州立大学; 康涅狄格大学; 威廉与玛丽学院; 巴塞罗那超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GroupMask,通过超网络生成层自适应分组稀疏比,在半结构化剪枝中优于均匀N:M模式,在LLaMA-2-7B上降低困惑度并提升零样本准确率。
AI 中文摘要
半结构化剪枝在压缩大型语言模型(LLM)的同时保持规则的稀疏结构,但主流的N:M模式在每个层中都固定了相同的局部稀疏比。层自适应稀疏性分配改善了非结构化剪枝,但据报道在N:M稀疏性下效果较差,这引出一个问题:自适应分配对半结构化剪枝总体而言价值有限,还是仅在细粒度N:M模式下价值有限?我们通过分组稀疏性来研究这一问题,分组稀疏性将每个权重矩阵划分为规则的分组,整体保留或剪除每个分组,并允许每个层的稀疏比在全局预算下变化。我们提出了GroupMask,它使用一个轻量级超网络生成所有层的分组选择器,通过Gumbel-Sigmoid参数化和直通估计器对其松弛化,并通过稀疏性预算正则化和自蒸馏进行学习,同时保持预训练权重冻结。在LLaMA-2-7B上,在50%稀疏度且分组大小为$1\ imes256$时,与均匀逐层稀疏比相比,学习的层自适应分配将WikiText-2困惑度从10.02降至8.30,并将平均零样本准确率从0.455提升至0.496。在五个LLaMA和Qwen模型上,与评估的基线相比,GroupMask在LLaMA-2-7B上取得了最低的WikiText-2困惑度,并在使用Alpaca校准的情况下取得了最高的平均零样本准确率。我们的代码可在https URL获取。
英文摘要
Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, leaving open whether adaptive allocation is of limited value for semi-structured pruning in general or only under the fine-grained N:M pattern. We examine this question with group-level sparsity, which partitions each weight matrix into regular groups, retains or prunes each group as a whole, and allows each layer's sparsity ratio to vary under a global budget. We propose GroupMask, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen. On LLaMA-2-7B at 50% sparsity with the same $1\times256$ group size, learned layer-adaptive allocation reduces WikiText-2 perplexity from 10.02 to 8.30 and raises the average zero-shot accuracy from 0.455 to 0.496 relative to a uniform per-layer ratio. GroupMask obtains the lowest WikiText-2 perplexity on LLaMA-2-7B and the highest average zero-shot accuracy with Alpaca calibration among the evaluated baselines on five LLaMA and Qwen models. Our code is available at https://github.com/ZhengaoLi/GroupMask.