通过搜索内核协同设计实现可实现的N:M稀疏Transformer推理
Realizable N:M Sparse Transformer Inference via Search-Kernel Co-Design
浏览论文内容
中文总结 AI 辅助
研究针对视觉Transformer推理延迟高的问题,提出硬件-软件协同设计框架。硬件设计MD-SpMM内核,软件进行分层稀疏搜索。在多模型和GPU平台实验,实现超2.2倍延迟加速,保持精度并在同延迟约束下有更好表现。
中文摘要 AI 辅助
视觉Transformer(ViT)虽精度高,但推理延迟大。半结构化N:M稀疏性可降低算术成本,却难在现代GPU上实现成比例的端到端加速,因部署延迟不仅取决于算术减少,还受稀疏性下的执行规则性和硬件调度影响。为此,提出N:M稀疏ViT推理的硬件-软件协同设计框架。硬件上设计MD-SpMM内核,软件上进行分层稀疏搜索。实验表明该框架在保持精度的同时,延迟加速超2.2倍。
英文摘要
Vision Transformers (ViTs) achieve strong accuracy but incur high inference latency. Semi-structured N:M sparsity can reduce arithmetic cost, yet its theoretical savings often fail to translate into proportional end-to-end speedups on modern GPUs. This mismatch arises because deployment latency depends not only on arithmetic reduction but also on execution regularity and hardware scheduling under sparsity. Achieving practical acceleration, therefore, requires coordinated design across sparse execution and sparsity configuration. To this end, we propose a hardware-software co-design framework for N:M sparse ViT inference. On the hardware side, we design MD-SpMM, an N:M sparse CUDA kernel that reorganizes sparse GEMM into micro-dense, Tensor-Core-aligned dataflow and uses inference-aware adaptive parallelism to sustain utilization. On the software side, we perform layer-wise sparsity search under explicit end-to-end latency budgets using a three-stage heuristic search with constraint relaxation to avoid premature convergence and enable deployment-aware sparsity allocation. Experiments on multiple ViT/Swin models and GPU platforms show that the framework achieves over 2.2x latency speedup while maintaining comparable accuracy and delivering superior accuracy under the same latency constraint. The source code is publicly available at https://github.com/liuganhuo/realizable-nm-sparse-transformer.