arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04093cs.LGcs.DC

超越掩码稀疏性:SNACK 在 GPU 上实现真正的稀疏神经网络

Beyond Masked Sparsity: SNACK Enables Truly Sparse Neural Networks on GPU

Jafar Badour, Maurice van Keulen, Elena Mocanu

首次发表
浏览论文内容

中文总结 AI 辅助

SNACK 提出一种真正的稀疏 GPU 层,仅存储和计算非零连接,通过自定义 COO 格式 SpMM 内核,在 90% 稀疏度下实现训练和推理加速,并显著降低内存和能耗。

中文摘要 AI 辅助

深度神经网络的参数数量持续增长,导致在 GPU 上的训练和推理成本上升。稀疏神经网络和动态稀疏训练(DST)有望降低这些成本,但大多数实现依赖于对稠密张量的二进制掩码,并且仅能恢复理论计算、内存或能量节省的一小部分。我们提出了 SNACK,一种真正的稀疏 GPU 层,它仅存储和计算非零连接。SNACK 提供了简单的 PyTorch API,用于在完全稀疏的范式下重构连接和反向传播梯度,并提供了 SNACK-COO,一个自定义的 COO 格式 SpMM CUDA 内核,其批量到流式多处理器映射针对大模型训练和单流推理中典型的小批量、高稀疏度场景进行了调优。在内核层面,SNACK 比掩码稠密基线(Dense+Mask)快最多 7 倍,并且在 95% 稀疏度下与 cuSPARSE、Sputnik 和 FlashSparse 相当。在 90% 稀疏度下,单个 SNACK 层相比 Dense+Mask 和完全稠密层,训练加速分别达到 8 倍和 3.7 倍,推理加速分别达到 4 倍和 2 倍,同时比稠密层少使用 72% 的内存,并显著降低能耗。端到端地,SNACK 在 99% 稀疏度下,相比 Dense+Mask,将 GPT-2 的峰值训练内存减少最多 40%,并将图式推理延迟降低 4.8 倍。

英文摘要

Deep neural networks continue to grow in parameter count, driving up training and inference cost on GPUs. Sparse neural networks and Dynamic Sparse Training (DST) promise to reduce these costs, but most implementations rely on binary masks over dense tensors and recover little of the theoretical compute, memory, or energy savings. We propose SNACK, a truly sparse GPU layer that stores and computes only non-zero connections. SNACK exposes a simple PyTorch API for restructuring connections and backpropagating gradients entirely in the sparse paradigm, and ships SNACK-COO, a custom COO-format SpMM CUDA kernel with a batch-to-Streaming-Multiprocessor mapping tuned for the small-batch, high-sparsity regime typical of large-model training and single-stream inference. At the kernel level, SNACK is up to 7x faster than the masked dense baseline (Dense+Mask) and competitive with cuSPARSE, Sputnik, and FlashSparse at 95% sparsity. At 90% sparsity, a single SNACK layer accelerates training by 8x and 3.7x, and inference by 4x and 2x, over Dense+Mask and fully dense layers, respectively, while using 72% less memory than dense and substantially less energy. End-to-end, SNACK reduces GPT-2 peak training memory by up to 40% and graph-style inference latency by 4.8x over Dense+Mask at 99% sparsity.

发表机构

  • University of Twente(特文特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑