在GPU CUDA及Tensor Core上实现支持多精度的内存高效Im2win卷积
Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores
浏览论文内容
中文总结 AI 辅助
该研究扩展im2win卷积范式,在GPU CUDA核心和Tensor Core上实现多精度支持,经12个CNN基准测试,其性能优于cuDNN等现有方法,内存开销显著降低,成为统一高性能卷积框架。
中文摘要 AI 辅助
卷积是深度神经网络的主要计算瓶颈,其效率取决于算法与GPU硬件的紧密集成。现有GPU卷积方法存在内存开销大、缓存利用率低、在不同卷积核尺寸上效果有限或数值不稳定等问题。本研究扩展了im2win范式——一种适用于所有卷积核尺寸、内存高效且内存访问连续的通用卷积方法——使其能在CUDA核心上以全精度运行,在Tensor Core上以半精度运行。通过引入新的卷积核设计与优化,如锯齿状内存访问和异步数据移动,im2win可高效利用硬件加速的半精度矩阵乘累加运算。在12个CNN基准测试中,im2win的TFLOPS最高比其CUDA核心实现高2.8倍,比cuDNN高1.4倍,比基于GEMM与cuBLAS的卷积高6.4倍,同时分别仅使用了后者53%和35%的内存。这些结果确立了im2win为适用于现代GPU架构的统一高性能卷积框架。
英文摘要
Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware. Existing GPU convolution methods suffer from large memory overhead, poor cache utilization, limited effectiveness across kernel sizes, or numerical instability. This work extends the im2win paradigm -- a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes -- to run efficiently in full precision on CUDA cores and half precision on tensor cores. By introducing new kernel designs and optimizations such as zig-zag memory access and asynchronous data movement, im2win efficiently exploits hardware-accelerated half-precision matrix multiply-accumulate operations. Across twelve CNN benchmarks, im2win achieves up to 2.8x higher TFLOPS than its CUDA core implementation, 1.4x higher than cuDNN, and 6.4x higher than GEMM-based convolution with cuBLAS, while using as little as 53% and 35% of their memory, respectively. These results establish im2win as a unified, high-performance convolution framework for modern GPU architectures.
发表机构
- Nanchang Hangkong University(南昌航空大学)
- EEO Tech(EEO科技)
- University of Washington(华盛顿大学)
- AWS(亚马逊云计算服务)
机构由 AI 辅助整理,请以论文原文为准。