UnionSparse:面向边缘设备低比特稀疏大语言模型推理的索引高效稀疏性框架
UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge
浏览论文内容
中文总结 AI 辅助
该研究针对边缘低比特稀疏LLM推理中索引流量等SpMM瓶颈,提出UnionSparse框架,结合IE-BME与LSPD内核,在W4A4量化及30%-70%稀疏度下实现优于现有方案的性能。
中文摘要 AI 辅助
边缘设备上的大语言模型(LLM)推理结合稀疏性与低比特量化,以满足设备内存、延迟和功耗限制。然而,量化会压缩权重有效载荷,却未按比例减少稀疏元数据,导致索引流量和非零元素提取成为稀疏矩阵乘法(SpMM)的关键瓶颈。我们提出有效载荷与元数据比率(PMR),证明提升PMR可提高解码阶段的有效计算强度。我们推出UnionSparse,这是一个索引高效框架,结合了索引高效位图编码(IE-BME)与采用低比特共享内存并行解码(LSPD)的SpMM内核。IE-BME分摊元数据开销并使稀疏遍历与碎片组装对齐,而LSPD可提升小批量执行效率。在W4A4量化和30%至70%稀疏度条件下,UnionSparse的性能分别超过FlashLLM和SpInfer 2.30倍、1.43倍,超过CUTLASS和cuBLAS Tensor Core 1.56倍、3.46倍。这些结果确立了有效载荷提取效率是边缘GPU上低比特稀疏推理的首要关注点。源代码可在以下网址获取:this https URL。
英文摘要
Edge LLM inference combines sparsity and low-bit quantization to meet device memory, latency, and power limits. Yet quantization shrinks weight payloads without proportionally reducing sparse metadata, so index traffic and nonzero extraction become critical SpMM bottlenecks. We introduce the Payload-to-Metadata Ratio (PMR) and show that improving PMR raises effective compute intensity in decoding. We present UnionSparse, an index-efficient framework that combines Index-Efficient Bitmap Encoding (IE-BME) with a SpMM kernel using Low-Bit Shared-Memory Parallel Decoding (LSPD). IE-BME amortizes metadata and aligns sparse traversal with fragment assembly, while LSPD improves small-batch execution. Under W4A4 quantization and 30%--70% sparsity, UnionSparse outperforms FlashLLM and SpInfer by 2.30x and 1.43x, and CUTLASS and cuBLAS Tensor Core by 1.56x and 3.46x, respectively. These results establish payload-extraction efficiency as a first-order concern for low-bit sparse inference on edge GPUs. Source code is available at: https://github.com/Victor-Alen/UnionSparse.