arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12327cs.LGcs.CL

SparseDecoding:面向准确高效的大语言模型推理的解码感知剪枝方法

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

Qitong Wang, Xinwei Niu, Mingluo Su, Shanwei Zhao, Shiai Zhu, Huan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

SparseDecoding是一种解码感知的LLM剪枝框架,通过对齐剪枝目标与解码激活、优化SpMV核,在长文本生成任务中优于标准方法,在A100 GPU上实现最高1.48倍解码加速。

中文摘要 AI 辅助

大语言模型(LLM)推理的解码阶段具有内存密集型特性,会导致显著的延迟。基于Hessian的逐层无训练网络剪枝方法是解决该问题的重要方案,因为剪枝可减少解码时从内存读取的非零参数数量。然而,该类典型方法使用预先收集的自然序列计算Hessian,而模型在解码时输入的是自生成的token,两种序列间存在分布偏移,导致自然序列上计算的Hessian与生成序列上的不同。我们发现这种差异会使生成时的激活分布偏离剪枝所用的分布,进而损害剪枝后模型的性能。此外,多数能带来实际加速的现有LLM剪枝方法主要针对稀疏矩阵-矩阵(SpMM)乘法,对主导解码过程的稀疏矩阵-向量(SpMV)操作支持有限。为解决这些问题,我们提出SparseDecoding,这是一种专为准确高效的LLM解码设计的原则性解码感知剪枝框架。具体而言,在算法层面,SparseDecoding从密集模型自回归生成(排除预填充阶段)期间收集的逐层激活中构建校准矩阵,使剪枝目标与解码激活对齐;在系统层面,我们开发了一种带位掩码索引和固定步长遍历的优化N:M稀疏矩阵-向量核。在代表性LLM(Llama-3.1-8B、Llama-3.3-70B、Qwen3-14B / 32B)上的大量实验结果表明,我们的方法在长文本生成基准测试中始终优于标准固定文本校准方法,且在A100 GPU上实现了最高达1.48倍的端到端挂钟解码加速。

英文摘要

The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding. Nevertheless, typical methods in this line compute the Hessian using pre-collected natural sequences, whereas the model is fed self-generated tokens during decoding, creating a distribution shift between the two sequences. The Hessian calculated on the natural sequence is different from that calculated on the generated sequence. We observe that this discrepancy causes the activation distribution during generation to deviate from that used for pruning, further hurting the pruned model performance. Moreover, most existing LLM pruning methods that bring actual speedup primarily target the sparse matrix-matrix (SpMM) multiplication, providing limited support for the sparse matrix-vector (SpMV) operations, which dominate decoding. To solve these problems, we introduce SparseDecoding, a principled decoding-aware pruning framework tailored for accurate and efficient LLM decoding. Specifically, at the algorithmic axis, SparseDecoding constructs calibration matrices from layer-wise activations collected during the dense-model autoregressive generation, excluding prefill, thereby aligning the pruning objective with the decoding activations. At the system axis, we develop an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal. Substantial empirical results on representative LLMs (Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B / 32B) demonstrate that our method consistently outperforms standard fixed-text calibration on the long-form generation benchmarks while achieving up to 1.48x end-to-end wall-clock decoding speedup on A100 GPUs.

发表机构

  • Westlake University(西湖大学)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

↑