发表机构
Yale University; University of Southern California(耶鲁大学; 南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对双稀疏LLM推理中spMspV workload处理低效问题,提出协同设计的Celty稀疏格式、GPU内核与SIMT微架构,实现最高5.3倍加速。
AI 中文摘要
大语言模型(LLMs)日益依赖稀疏性降低推理成本,但多数现有工作仅针对单一稀疏源(权重或激活),且针对批量多用户推理优化。双稀疏性结合非结构化权重剪枝与运行时激活稀疏性,在单用户解码中实现模型规模、精度与延迟的良好权衡,但其表现为稀疏矩阵-稀疏向量(spMspV) workload,现有GPU内核处理效果不佳。本文提出Celty,一种协同设计的稀疏格式、GPU内核与SIMT微架构,用于LLM推理中高效处理spMspV。内核层面,Celty引入游程压缩CSC(RLC-CSC)格式,支持压缩权重列的向量化加载,利用两种稀疏性跳过不必要的内存访问,采用共享内存进行分散的部分乘积累加。微架构层面,Celty稀疏SIMT核心集成流水线RLC解码器,消除软件级索引重建,将本地寄存器文件重新用于无冲突累加,直接基于相同RLC-CSC格式操作,无需数据布局变更。Celty GPU内核相比cuBLAS最高实现2.8倍加速,相比Flash-LLM最高实现2.4倍加速;采用稀疏SIMT核心后,在70%双稀疏度下相比cuBLAS最高实现5.3倍加速。
英文摘要
Large Language Models (LLMs) increasingly rely on sparsity to cut inference cost, but most prior work exploits a single sparsity source and targets batched multi-user inference. Dual-sparsity, which pairs unstructured weight pruning with runtime activation sparsity, offers a compelling size-accuracy-latency tradeoff for single-user decoding, but forms a Sparse Matrix-Sparse Vector (SpMSpV) workload that existing GPU kernels handle poorly. We propose Celty, a co-designed sparse format, GPU kernel, and SIMT microarchitecture for SpMSpV in LLM inference. Celty's Run-Length Compressed CSC (RLC-CSC) format enables vectorized loading of compressed weight columns and exploits both sparsity sources to skip memory accesses, accumulating partial products in shared memory. The Celty Sparse SIMT Core then adds a pipelined RLC decoder that removes software index reconstruction and repurposes local register files for conflict-free accumulation, operating on the same compressed representation. The kernel alone achieves up to 2.8x over cuBLAS; with the Sparse SIMT Core, speedup reaches 5.3x over cuBLAS at 70% dual-sparsity.
CommentsICCAD 2026. The code is available on Github at https://github.com/RuokaiYin/Celty