发表机构
Blubridge AI(布鲁布里奇人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Nova是一款端到端MLIR编译器,通过整图优化等技术提升深度学习模型性能,在RTX 3060上的TF32矩阵乘法、模型吞吐量和内存占用等指标上优于PyTorch和XLA,可训练更大规模模型且避免OOM故障。
AI 中文摘要
大规模深度学习模型的性能高度依赖于高层数学运算到底层物理硬件的映射效率。尽管高层张量框架为模型设计提供了灵活的抽象,但它们的即时执行模型天生缺乏最大化物理硬件原生利用率所需的整图可见性,以及对硬件和内存的细粒度控制。为了弥合这一差距,我们设计了Nova,一款自动化端到端JIT编译器,其核心目标是实现对硬件映射的绝对控制:跨运算边界融合运算、优化复杂内存层次结构、将执行调优至寄存器级别。通过捕获即时执行并将前向与反向传播统一为单一值语义方言,Nova实现了激进的整图优化。随后,它利用分析配置器根据算术强度确定性推导最优执行调度,将搜索时间降至零。在结构哈希运行时的支持下,Nova直接从计算结构合成细粒度内核。在RTX 3060上的评估显示,在大多数形状的TF32矩阵乘法中,Nova的性能与cuBLAS和XLA相当或略有超出,同时保持严格的<5e-4相对误差。在模型层面,对于一个4200万参数的模型,Nova的吞吐量比PyTorch高10.6%,比XLA高4.4%,且不损害数值保真度。至关重要的是,与PyTorch相比,Nova将内存占用最多降低29%,从而在同一块12GB消费级GPU上,PyTorch会出现内存不足(OOM)故障的情况下,Nova能以17900 token/s的速度成功训练一个1.44亿参数的模型。
英文摘要
The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. While high-level tensor frameworks provide flexible abstractions, their execution models inherently lack the whole-graph visibility required to maximize hardware utilization, often forcing a reliance on opaque, hand-written kernel libraries for complex operations like Attention. To bridge this gap, we present the next iteration of Nova, an automated end-to-end JIT compiler that achieves absolute control over hardware mapping by synthesizing fine-grained kernels directly from the computation's structure. In this work, we extend Nova's compilation pipeline to natively support full Transformer architectures. By capturing eager executions and unifying forward and backward passes into a single value-semantic dialect, Nova unlocks aggressive whole-graph optimizations. Rather than relying on rigid, pre-compiled library calls, Nova focuses on extensive cross-operator fusions, collapsing complex causal attention sub-graphs, element-wise operations, and memory-bound normalizations directly into single fused kernels to drastically reduce global memory roundtrips. In our evaluations training a full GPT-2 architecture on Ada 6000 GPUs, Nova demonstrates superior end-to-end throughput, averaging 441K tokens/second compared to 406K for our own eager execution and 405K for torch.compile. By drastically reducing memory-bound overheads through compiler-native fusion, Nova enables efficient full LLM compilation on modern hardware while strictly maintaining numerical parity.