释放等式饱和在张量程序超级优化中的力量
Unleashing the Power of Equality Saturation for Tensor Program Superoptimization
浏览论文内容
中文总结 AI 辅助
EqiForge利用等式饱和统一IR,联合优化张量程序的高层代数与低层执行,实现1.32倍平均加速,并发现超越FlashAttention的注意力内核。
中文摘要 AI 辅助
高效的GPU张量程序实现通常需要联合优化高层代数公式和低层执行策略。然而,随着变换跨算子组合,所产生的搜索空间迅速增长,使得联合优化难以扩展。我们提出了EqiForge,一个基于等式饱和的张量程序超级优化器。其统一IR在单一表达式语言中表示高层张量表达式和分块计算。通过组合等式规则,EqiForge直接从张量表达式推导出融合实现,如FlashAttention风格的内核。早期压缩在完成之前剪除冗余的部分程序,而子图组合将搜索扩展到更大的图。在张量程序基准测试中,EqiForge相对于每个配置的最快可用基线实现了1.32倍的几何平均加速和最大2.74倍的加速。其注意力内核在解码方面比FlashAttention快高达1.87倍,并在预填充方面接近其性能。EqiForge还发现了新的实现,在各种Transformer层上优于此http URL,包括QK归一化MLA(3.16倍)和mHC(5.84倍)。
英文摘要
Efficient GPU implementations of tensor programs often require joint optimization of high-level algebraic formulations and low-level execution strategies. However, the resulting search space grows rapidly as transformations combine across operators, making joint optimization difficult to scale. We present EqiForge, a tensor program superoptimizer based on equality saturation. Its unified IR represents high-level tensor expressions and tiled computations in a single expression language. By composing equality rules, EqiForge derives fused implementations such as FlashAttention-style kernels directly from tensor expressions. Early compaction prunes redundant partial programs before completion, while subgraph composition extends the search to larger graphs. Across tensor-program benchmarks, EqiForge achieves a geometric mean speedup of 1.32x and a maximum of 2.74x over the fastest available baseline per configuration. Its attention kernels outperform FlashAttention by up to 1.87x in decode and approach its performance in prefill. EqiForge also discovers new implementations that outperform torch.compile on various Transformer layers, including QK-normalized MLA (3.16x) and mHC (5.84x).
发表机构
- Zhejiang University(浙江大学)
- Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究所)
机构由 AI 辅助整理,请以论文原文为准。