TriSP:用于大语言模型的三信号结构化剪枝
TriSP: Tri-Signal Structured Pruning for Large Language Models
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型部署受参数成本限制问题,提出TriSP方法,结合权重幅度、激活范数与梯度敏感性产生通道级分数,联合自适应预算分配和低秩适应恢复,实现低困惑度、高零样本准确率及高推理吞吐量。
中文摘要 AI 辅助
大语言模型在各种任务中表现出色,但参数的内存和计算成本限制了其部署。结构化剪枝通过移除注意力头和多层感知器神经元等结构来解决此问题。现有方法存在依赖基于梯度的重要性估计(内存消耗大)或基于激活的统计代理(不直接衡量移除对损失的影响)等问题,且重要性标准与剪枝后恢复策略的相互作用未被系统研究。我们提出TriSP,它通过几何均值结合由激活范数缩放的权重幅度和一阶梯度敏感性,产生捕捉结构和损失敏感性信号的通道级分数。结合自适应每层预算分配和低秩适应恢复,TriSP在所有测试配置中实现了最低困惑度和最高零样本准确率,在对LLaMA - 7B进行20%剪枝时达到6.80的WikiText - 2困惑度。在50%剪枝时推理吞吐量提高82%,同时保持竞争力。
英文摘要
Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their parameters. Structured pruning addresses this by removing entire structures such as attention heads and Multi-Layer Perceptron (MLP) neurons to produce smaller dense models that run efficiently on standard hardware. However, existing methods rely on either gradient-based importance estimation, which is memory-prohibitive, or activation-based statistical proxies, which do not directly measure the effect of removal on the loss. Furthermore, the interaction between the importance criterion and the post-pruning recovery strategy has not been systematically studied. We propose TriSP (Tri-Signal Structured Pruning), an importance metric that combines weight magnitude scaled by activation norm with first-order gradient sensitivity via a geometric mean, producing a channel-level score that captures both structural and loss-sensitivity signals. Combined with adaptive per-layer budget allocation and low-rank adaptation (LoRA) recovery, TriSP achieves the lowest perplexity and highest zero-shot accuracy across all tested configurations, reaching 6.80 WikiText-2 perplexity at 20% pruning on LLaMA-7B. Inference throughput improves by 82% at 50% pruning, while still maintaining competitive performance.
发表机构
- National School of Artificial Intelligence (ENSIA)(国家人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。