arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ComFuse:面向现代GPU架构融合复杂内存密集型子图与计算密集型内核

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures

Di Mu, Tengyuan Jin, Zhenkun Wang, Jialin Yang, Yusen Li, Mian Huo, Shusong Guo, Gang Wang, Xiaoguang Liu

arXiv 2608.03537首次发表:更新:

AI 中文总结

本文提出的自动化GPU编译系统ComFuse,通过新颖的算子融合策略处理异构深度学习计算图,支持B2BGEMM模式融合,生成的融合内核性能优于TorchInductor,且融合模式更灵活。

AI 中文摘要

现代深度学习 workload 越来越多地包含异构计算图,将计算密集型算子与内存密集型子图相结合。现有深度学习编译器通常单独优化这些算子类别,形成僵化的融合边界,限制了跨算子优化和片上数据复用。我们观察到下游内存密集型操作可与计算密集型算子并发执行,使其执行可隐藏在计算之后;但自动利用该机会带来了新的编译挑战。本文提出 ComFuse,这是一种自动化 GPU 编译系统,采用新颖的算子融合策略,为包含计算密集型算子、依赖关系丰富的内存密集型元素归约子图的复杂图结构生成高性能内核。ComFuse 还支持背对背 GEMM(B2BGEMM)模式的融合,将其适用性扩展到更复杂的计算-内存交互模式。此外,它自动将高级张量子程序 lower 为优化的融合内核,减少了手动内核工程的需求。实验结果表明,ComFuse 生成的融合内核在归一化后工作负载及各种复杂计算场景中,性能优于 TorchInductor 生成的内核,同时支持更灵活的融合模式。

英文摘要

Modern deep learning workloads increasingly comprise heterogeneous computation graphs that combine compute-intensive operators with memory-intensive subgraphs. Existing deep learning compilers typically optimize these operator classes separately, creating rigid fusion boundaries that limit cross-operator optimization and on-chip data reuse. We observe that downstream memory-intensive operations can execute concurrently with compute-intensive operators, allowing their execution to be hidden behind computation; however, automatically exploiting this opportunity poses new compilation challenges. In this paper, we present ComFuse, an automated GPU compilation system that employs a novel operator fusion strategy to generate high-performance kernels for complex graph structures comprising compute-intensive operators and dependency-rich, memory-intensive elementwise-reduction subgraphs. ComFuse further supports the fusion of back-to-back GEMM (B2BGEMM) patterns, extending its applicability to more complex compute-memory interaction patterns. Additionally, it automatically lowers high-level tensor subprograms into optimized fused kernels, reducing the need for manual kernel engineering. Experimental results show that the fused kernels generated by ComFuse outperform those produced by TorchInductor across post-norm workloads and various complex computation scenarios, while supporting more flexible fusion patterns.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑