发表机构
National University of Singapore; Zhejiang University(新加坡国立大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型推理成本挑战,提出含稀疏张量核层等的三层矩阵存储格式,设计联合稀疏张量核与CUDA核的SpMM内核,实现了在现代GPU上超越密集矩阵乘法,取得显著加速。
AI 中文摘要
随着大语言模型(LLMs)部署的增加,其推理成本成为关键挑战。将稀疏性引入权重矩阵的剪枝技术可加速推理,但维持模型质量通常将剪枝限制在适度的非结构化稀疏性(约50%)。在此稀疏水平下,现有的稀疏矩阵乘法(SpMM)GPU内核均无法超越其密集对应物。本文提出一种针对适度稀疏性LLMs的高效GPU推理方法。我们提出一种三层矩阵存储格式,包括:(i)一个稀疏张量核层以加速SpMM;(ii)一个插槽填充层用于矩阵压缩并支持低成本片上解码;(iii)一个轻量级残差层确保正确的SpMM计算。基于此格式,我们设计了一个联合利用稀疏张量核和CUDA核的SpMM内核。评估表明,我们的工作首次在配备高带宽内存(HBM)的现代GPU上超越了密集矩阵乘法。它比SpInfer实现了高达1.64倍的内核级加速,比FlashLLM实现了高达1.41倍的端到端加速。我们的源代码:此https URL。
英文摘要
With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can accelerate inference. However, maintaining model quality typically limits pruning to moderate unstructured sparsity (around 50\%). At these sparsity levels, none of the existing GPU kernels for sparse matrix multiplication (SpMM) can outperform their dense counterparts. This paper proposes an efficient GPU inference method for LLMs with moderate sparsity. We propose a three-layer matrix storage format comprising: (i) a Sparse-TC layer enabling sparse tensor cores to accelerate SpMM; (ii) a Slot-Filling layer using parallel differential distance for matrix compression while supporting low-cost on-chip decoding; (iii) a lightweight Residual Layer ensuring correct SpMM computation. Building on this format, we design a SpMM kernel that jointly utilizes sparse tensor cores and CUDA cores. This design enables an efficient execution pipeline and overlaps on-chip computation with memory access. Evaluations show that our work is the first to outperform dense matrix multiplication on modern GPUs equipped with high-bandwidth memory (HBM). It achieves up to 1.64x kernel-level speedup over SpInfer (EuroSys'25, Best paper) and up to 1.41x end-to-end speedups over FlashLLM (VLDB'24). Our source code: https://github.com/moui0/cudac.
CommentsDAC 2026