发表机构
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi; Indian Institute of Technology Delhi; Microsoft Research (MSR) India(德里印度理工学院亚迪人工智能学院; 德里印度理工学院; 微软研究院印度分部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLMs结构化剪枝的贪心启发式算法缺陷,提出SNIPER两阶段剪枝框架,引入CRAFT量化预算保真度,在多架构多任务上较现有剪枝器实现性能与稳定性提升,泛化性强。
AI 中文摘要
结构化剪枝是压缩大语言模型(LLMs)的有效方法,但现有方法严重依赖贪心启发式算法,会产生短视决策,且常无法精准满足目标压缩预算。我们提出SNIPER,一种两阶段结构化剪枝框架,该框架在粗粒度组件上求解背包优化问题,以生成基于固定重要性估计的条件最优参数分配,随后通过细粒度剪枝阶段满足严格的预算约束。我们引入压缩比率 adherence 因子(CRAFT)来量化预算保真度,结果显示,现有剪枝器与目标压缩比率的偏差最高达33%,而SNIPER的CRAFT得分为0.98,实现了近乎精确的 adherence。在五个领域的18项任务组成的集合上,对四种不同架构的评估表明,与六种最先进的剪枝器相比,SNIPER在平均性能保持和任务级稳定性方面实现了持续改进。在所有剪枝配置中,SNIPER的平均排名为1.25,表明其具有强大的跨架构泛化能力和出色的可靠性。
英文摘要
Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets. We present SNIPER, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to yield conditionally optimal parameter allocations with respect to fixed importance estimates, followed by a fine-grained pruning stage to meet strict budget constraints. We introduce the Compression Ratio Adherence Factor (CRAFT) to quantify budget fidelity, showing that while existing pruners deviate from target compression ratios by up to 33%, SNIPER achieves near-exact adherence with a CRAFT score of 0.98. Evaluations across four diverse architectures over a set of 18 tasks spanning five domains demonstrate SNIPER's consistent improvements in average performance retention and task-level stability over six state-of-the-art pruners. Across all pruning configurations, SNIPER achieves an excellent mean rank of 1.25, indicating its robust cross-architectural generalizability and excellent reliability.
Comments29 pages, 5 figures, 17 tables