AI 中文总结
本文提出面向动态GPU工作负载的双视角分块SpMSpV框架DB-SpMSpV,集成至DB-BFS和DB-Decoding后,在A100和RTX 4090上相比现有方法实现了显著加速。
AI 中文摘要
稀疏矩阵-稀疏向量乘法(SpMSpV)是图遍历、稀疏线性代数和稀疏模型推理中的核心基础操作。其输入向量通常具有动态稀疏性,因此最佳GPU执行路径取决于全局稀疏度和局部向量块分布。现有GPU SpMSpV方法常将存储布局、推/拉遍历与内核绑定,导致在无额外存储或调度开销时难以实现细粒度适配。本文提出DB-SpMSpV,一种面向动态GPU工作负载的双视角分块SpMSpV框架。DB-SpMSpV将矩阵划分为固定大小的2D块,在高层维护块级CSR/CSC视角,复用单个低层块载荷以支持行驱动的拉取和列驱动的推送。运行时,它根据输入块稀疏度选择全局遍历路径,从局部矩阵/向量块结构中选择块微内核,并通过负载均衡、异步预取和分层写回减少不规则内存访问、写回冲突和负载不均衡。我们进一步将该框架集成到DB-BFS和DB-Decoding中。我们在NVIDIA A100和RTX 4090上使用SuiteSparse矩阵、对称图和三个开源大语言模型(LLM)对DB-SpMSpV进行评估。在各类输入稀疏度下,DB-SpMSpV在A100上相比cuSPARSE实现了5.48倍至64.34倍的平均加速,相比TileSpMSpV实现了2.36倍至14.01倍的加速,在RTX 4090上获得了类似的增益。DB-BFS在A100上相比TileBFS将端到端图遍历平均提升2.66倍,在RTX 4090上提升3.60倍,而DB-Decoding将单token线性层加速最高达4.50倍。
英文摘要
Sparse Matrix-Sparse Vector Multiplication (SpMSpV) is a core primitive in graph traversal, sparse linear algebra, and sparse model inference. Its input vector is often dynamically sparse, so the best GPU execution path depends on both global sparsity and the local vector-block distribution. Existing GPU SpMSpV methods often bind storage layouts, push/pull traversal, and kernels together, making fine-grained adaptation difficult without extra storage or scheduling overhead. This paper presents DB-SpMSpV, a dual-view blocked SpMSpV framework for dynamic GPU workloads. DB-SpMSpV partitions the matrix into fixed-size 2D blocks, maintains block-level CSR/CSC views at the high level, and reuses a single low-level block payload to support both row-driven pull and column-driven push. At runtime, it selects the global traversal path based on input block sparsity, chooses block microkernels from the local matrix/vector block structure, and uses load balancing, asynchronous prefetching, and hierarchical writeback to reduce irregular memory accesses, writeback conflicts, and load imbalance. We further integrate the framework into DB-BFS and DB-Decoding. We evaluate DB-SpMSpV on NVIDIA A100 and RTX 4090 using SuiteSparse matrices, symmetric graphs, and three open-source LLMs. Across input sparsities, DB-SpMSpV achieves average speedups of 5.48$\times$--64.34$\times$ over cuSPARSE and 2.36$\times$--14.01$\times$ over TileSpMSpV on A100, with similar gains on RTX 4090. DB-BFS further improves end-to-end graph traversal by 2.66$\times$ over TileBFS on A100 and 3.60$\times$ on RTX 4090 on average, while DB-Decoding accelerates single-token linear layers by up to 4.50$\times$.
Comments11 pages, 10 figures, Accepted by ICPP 2026;