从可移植到高效:在Julia中自动调优硬件无关的GPU内核
Portable to Efficient: Auto-Tuning Hardware-Agnostic GPU Kernels in Julia
浏览论文内容
中文总结 AI 辅助
该研究针对硬件无关GPU内核执行效率低的问题,通过在Julia中集成自动调优、重建Kernel Tuner框架,在SVD任务上使内核性能提升3-7倍,为异构GPU开发提供了高效方案。
中文摘要 AI 辅助
传统上,GPU内核是在供应商特定的编程模型中开发和优化以实现高性能的,这导致软件难以在日益异构的计算系统中进行优化和适配。硬件无关的编程模型通过提高可移植性和可维护性,为GPU软件开发提供了更可持续的方法,但在不同架构上实现高效执行仍然具有挑战性。我们通过将自动调优集成到用Julia编写的硬件无关GPU内核中来解决这一挑战。我们用Julia支持重建了成熟的Kernel Tuner自动调优框架,该框架能够系统地探索针对NVIDIA、AMD、Intel和Apple GPU的硬件无关GPU内核的配置。我们在硬件无关的奇异值分解(SVD)上演示了该方法,该SVD在the this http URL线性代数库中实现。结果表明,自动调优对于在各种硬件上创建资源高效的硬件无关GPU内核至关重要。与中位参数配置相比,最优配置将内核性能提高了3倍至7倍,证明了调优对高效硬件利用的重大影响。
英文摘要
Traditionally, GPU kernels have been developed and optimized within vendor-specific programming models to achieve high performance, resulting in software that is difficult to optimize and adapt across increasingly heterogeneous computing systems. Hardware-agnostic programming models offer a more sustainable approach to GPU software development by improving portability and maintainability, but achieving efficient execution across diverse architectures remains challenging. We address this challenge by integrating auto-tuning into hardware-agnostic GPU kernels written in Julia. We rebuild the established Kernel Tuner auto-tuning framework with Julia support, enabling systematic exploration of kernel configurations for hardware-agnostic GPU kernels targeting NVIDIA, AMD, Intel, and Apple GPUs. We demonstrate this approach on hardware-agnostic singular value decomposition (SVD) as implemented in the NextLA.jl linear algebra library. The results show that auto-tuning is essential for creating resource-efficient hardware-agnostic GPU kernels across a variety of hardware. Optimal configurations improve kernel performance by a factor of 3x to 7x compared to median parameter configurations, demonstrating the substantial impact of tuning on efficient hardware utilization.