arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DLCB:面向动态深度学习的 Ahead-of-Time 编译

DLCB: Ahead-of-Time Compilation for Dynamic Deep Learning

Alexander Collins, Bin Fan, Evghenii Gaburov, William Brandon, Sean Lee, Hanfeng Chen, Vinod Grover

arXiv 2610.10547首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DLCB 是一款嵌入 PyTorch 的深度学习编译器,通过三层编译策略解决动态深度学习场景中 AOT 编译性能与 JIT 编译开销的矛盾,支持不同动态程度的张量形状编译。

AI 中文摘要

深度学习工作负载越来越多地部署在编译时张量形状无法完全确定的场景中:不同请求的批量大小各不相同,输入的序列长度存在差异,模型架构也支持多种空间分辨率。这种动态性造成了一种根本矛盾:提前编译(Ahead-of-Time)的 GPU 内核能提供峰值性能,但传统上需要完全静态的张量形状;而即时编译(Just-in-Time)虽支持动态形状,却要以运行时编译开销为代价。我们提出 DLCB(Deep Learning Compiler Backend,深度学习编译器后端),一款通过统一的三层编译策略解决该矛盾的深度学习编译器。对于张量形状完全静态的程序,DLCB 提前生成静态 CUDA C++ 内核;对于张量秩在编译时已知但维度大小动态的程序,DLCB 生成以这些维度为参数的 AOT 内核,编译一次即可跨多种形状执行;对于张量秩未知的程序,DLCB 回退到运行时 JIT 编译。针对动态形状张量的 AOT 编译,我们的方法从输入程序构建形状约束系统,并用求解器将其简化为两类:a)固定约束,在生成的内核中产生硬编码常量;b)未解决约束,成为启动时传递的内核参数,同时搭配解释型主机代码,用于检查约束对给定输入集是否成立并计算要传递的值。我们的编译器和运行时嵌入在 PyTorch 中,采用自动形状泛化传递,为 PyTorch 程序的受支持子集实现动态形状的 AOT 编译。

英文摘要

Deep learning workloads are increasingly deployed in settings where tensor shapes are not fully known at compile time. Batch sizes vary across requests, sequence lengths differ between inputs, and model architectures admit a range of spatial resolutions. This dynamism creates a fundamental tension: ahead-of-time compiled GPU kernels deliver peak performance but traditionally require fully static tensor shapes, while just-in-time compilation supports dynamic shapes at the cost of runtime compilation overhead. We present DLCB (Deep Learning Compiler Backend), a deep learning compiler that resolves this tension through a unified three-tier compilation strategy. For programs with fully static tensor shapes, DLCB generates static CUDA C++ kernels ahead of time. For programs whose tensor rank is statically known but whose dimension sizes are dynamic, DLCB generates AOT kernels parameterized by those dimensions, compiling once and executing across a range of shapes. For programs with tensors of unknown rank, DLCB falls back to JIT compilation at runtime. For AOT compilation with dynamic shape tensors, our approach builds a system of shape constraints from the input program, and uses a solver to reduce these to either a) fixed constraints that produce hard-coded constants in the generated kernel; b) unresolved constraints which become kernel parameters passed at launch time, along with interpreted host code to check that the constraint holds for a given set of inputs and to compute the values to pass. Our compiler and runtime are embedded in PyTorch, and use an automatic shape generalization pass to enable AOT compilation with dynamic shapes for supported subsets of PyTorch programs

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑