发表机构
Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型推理瓶颈,提出统一静态-动态剪枝框架SPDP,集成非结构化SP与输入自适应DP,设计新格式和内核,经评估其加速效果显著,推进推理效率-质量前沿,提升大规模LLM服务的吞吐量和性能功耗比。
AI 中文摘要
大语言模型(LLM)部署增加,凸显自回归解码的计算和内存瓶颈,权重剪枝是补救方法,但现有方法局限于静态剪枝(SP)或动态剪枝(DP)。本文提出SPDP,统一稀疏推理框架,集成非结构化SP与输入自适应DP以在GPU上高效进行LLM推理。它共同设计新的平铺列位图压缩(Tiled-CBC)格式和两个互补GPU内核。综合评估表明,SPDP比SpInfer等加速1.24倍至1.37倍,在高达25%更高稀疏度下匹配困惑度,推进推理效率-质量帕累托前沿。
英文摘要
The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost. Weight pruning offers a promising remedy, but existing methods remain confined to either static pruning (SP), which permanently removes redundant weights but lacks adaptivity, or dynamic pruning (DP), which adapts to input sparsity but introduces runtime irregularity. This paper presents SPDP, a unified sparse-inference framework that integrates unstructured SP with input-adaptive DP for efficient LLM inference on GPUs. SPDP co-designs a new Tiled-Column-wise Bitmap Compressed (Tiled-CBC) format and two complementary GPU kernels: (1) a CUDA-core spMspV kernel featuring Hybrid Activation-aware Dynamic Shared-Memory Bitmap Decoding (HAD-SMBD) for fine-grained, runtime activation skipping, and (2) a Tensor-Core SpMM kernel optimized for prefill computation. This joint format-kernel design harmonizes static and dynamic sparsity, maintaining bandwidth-efficient memory access and high compute intensity under both phases of LLM inference. Comprehensive evaluations on inference-optimized GPUs demonstrate that SPDP achieves 1.24x-1.37x average speedup (up to 2.51x) over state-of-the-art sparse frameworks such as SpInfer, while matching perplexity with up to 25% higher sparsity. SPDP advances the inference efficiency-quality Pareto frontier, showing that unified static-dynamic pruning can deliver substantial throughput and performance-per-watt improvements in large-scale LLM serving.
CommentsProceedings of the VLDB Endowment (PVLDB), Volume 19, Issue 11, 2026