发表机构
University of British Columbia(不列颠哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
F-WANDA是WANDA的改进后训练剪枝方法,通过Fisher重加权分配保留预算,在LLAMA-2-7B上实现更优MMLU指标且能耗仅为SPARSEGPT的1/3,处于性能与剪枝成本的帕累托前沿。
AI 中文摘要
一次性后训练剪枝是大语言模型(LLM)最节能的压缩策略,但现有方法要么牺牲性能(如WANDA)要么增加计算成本(如SPARSEGPT)。本文提出F-WANDA,它是WANDA的直接改进版本,按预激活的经验Fisher信息重新分配输出神经元的逐行保留预算。Fisher信号通过在WANDA已使用的同一校准语料上额外执行一次反向传播收集,无需更新权重。在LLAMA-2-7B模型、50%非结构化稀疏度设置下,F-WANDA在WikiText-2上的困惑度为6.85,与WANDA的流畅度相当,5次设置的MMLU指标较WANDA提升1.6个百分点,较SPARSEGPT提升1.1个百分点,而剪枝的 wall-clock 时间和能耗仅为SPARSEGPT的三分之一。该核心权衡无需额外校准数据或微调即可实现,使F-WANDA处于LLM压缩的性能与剪枝成本帕累托前沿。
英文摘要
One-shot post-training pruning is the most energy-frugal compression strategy for largelanguage models (LLMs), yet existing approaches trade either quality (WANDA) or compute cost (SPARSEGPT). We introduce F-WANDA, a drop-in modification of WANDA that reallocates the per-row keep budget across output neurons in proportion to the empirical Fisher information of the pre-activation. The Fisher signal is collected in a single additional backward pass over the same calibration corpus WANDA already uses; no weights are updated. On LLAMA-2-7B at 50 % unstructured sparsity, F-WANDA attains WikiText-2 perplexity of 6.85, matches WANDA fluency, and improves 5-shot MMLU by +1.6 pp over WANDA and +1.1 pp over SPARSEGPT, while incurring only one-third of SPARSEGPT pruning wall-clock and energy. The headline trade-off is achieved without extra calibration data or fine-tuning, placing F-WANDA on the Pareto frontier of quality versus pruning cost for sustainable LLM compression.