arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Rust中的GPU卸载:可移植、安全且快速

GPU Offload in Rust: Portable, Safe, and Fast

Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala, Alán Aspuru-Guzik, Johannes Doerfert

arXiv 2608.13759首次发表:更新:

AI 中文总结

本文提出一种原生构建于rustc和LLVM后端的零开销多厂商GPU编译框架,利用Rust特性解决GPU编程的安全与效率问题,经RAJAPerf评估其内核性能可与CUDA、HIP C++基线媲美。

AI 中文摘要

高性能GPU编程传统上需要在执行效率和内存安全性之间做出权衡。Rust通过其严格的所有权模型为主机CPU提供编译时内存安全保障,而将这些约束应用于大规模并行GPU执行环境时,此前要么需要使用厂商锁定的领域特定语言(DSL),要么要转向显式的不安全原始指针。本文提出了一种零开销、多厂商GPU编译框架,该框架原生构建于Rust编译器(rustc)和LLVM后端中。我们利用Rust丰富的类型系统、所有权系统和严格的别名保障(noalias),通过LLVM的Offload基础设施高效管理和优化数据传输。我们揭示了主机与设备目标之间跨厂商ABI降低不匹配的技术挑战,并引入了一种两遍编译流水线,能够安全处理手动和编译器生成的内存移动。在RAJAPerf上对我们的框架进行评估,结果表明,基于rustc的解决方案可为GPU内核生成具有竞争力的LLVM IR,与原生手动优化的CUDA和HIP C++基线相比,内核性能表现出色。

英文摘要

High-performance GPU programming has traditionally forced a compromise between execution efficiency and memory safety. While Rust guarantees compile-time memory safety for host CPUs via its strict ownership model, applying these constraints to massively parallel GPU execution environments has previously mandated either vendor-locked Domain-Specific Languages (DSLs) or escaping to explicit unsafe raw pointers. This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends. We leverage Rust's rich type system, ownership system, and strict aliasing guarantees (noalias) to efficiently manage and optimize data transfers through LLVM's Offload infrastructure. We expose the technical challenges of cross-vendor ABI lowering mismatches between Host and Device targets and introduce a two-pass compilation pipeline capable of safely handling both manual and compiler-generated memory movements. Evaluating our framework on RAJAPerf demonstrates that our rustc-based solution can generate competitive LLVM IR for GPU kernels, achieving a solid kernel performance against native, hand-optimized CUDA and HIP C++ baselines.

Comments13 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑