arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于MTIA的Triton:弥合定制AI加速器的编程模型差距

Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators

Haishan Zhu, Domi Yan, Michael Levesque-Dion, Changxu Zhang, Mitch Gamburg, Kirsten Lee, Giancarlo Colmenares, Aditya Bhagwat, Arnab De, Markus Le Roux, Victor Perez Carrasco, Xin Tong, Will Cromar, Simran Barnwal, Andrew Uderian, Sridhar Gopinath, Jan Szczepaniec, Daniel Neilson, Blaine Burton Rister, Jordan Fix, Jazlyn Li, Zejun Huang, Lite Ye, Nan Zhang, Xinchen Guo, Andiry Xu, Michael Roberts, Kunming Ho, Site Cao, Suryadev Sahadevan Rajesh, Tristan Trouwen, Mike Tsai, Jake Lee, Wayne Su, Yuhan Chen, Xiaolong Xie, David Eklov, Aaron Barnes, Max Bremer, Adam Belay, Shintaro Iwasaki, Roman Levenstein, Ajit Mathews

arXiv 2608.00325首次发表:更新:

AI 中文总结

本研究将Triton应用于Meta定制ML加速器MTIA-2i,开发了专属编译器后端与TorchInductor增强,其内核性能可媲美专家调优的C++实现,已部署至生产环境覆盖大量模型,证明Triton可弥合编程模型差距。

AI 中文摘要

机器学习 workload 的快速增长推动了定制加速器架构的普及。这些从头设计的加速器通常呈现出与GPU不同的编程模型。尽管超大规模企业和AI芯片初创企业在该领域持续创新,但实现广泛的算子覆盖以支持多样化模型仍是一项重大挑战。此外,易用的高级内核编程语言对模型和内核的快速迭代至关重要。Triton与TorchInductor一起解决了GPU上的这些问题,但其在具有不同编程模型的加速器上的适用性尚未得到证实。在本研究中,我们首次将Triton应用于Meta开发的定制ML加速器MTIA-2i的生产规模场景。为支持MTIA-2i,我们开发了针对该加速器的新编译器后端,对TorchInductor的代码生成引入了增强,并提出了最小的语言扩展以暴露MTIA特有的架构特性。我们证明,Triton-MTIA内核的性能可与专家调优的C++实现相媲美。借助这些开发效率的提升,我们成功将手动编写和Inductor生成的Triton内核部署到生产环境中,覆盖了约60种不同的模型类型,占这些模型层数的50%和非GEMM执行时间的47%。我们的结果提供了有力证据,表明像Triton这样的DSL可以弥合ML框架、内核与定制加速器之间的编程模型差距,从而实现大规模的快速创新和高效部署。

英文摘要

The rapid growth in machine learning workloads has fueled the proliferation of custom accelerator architectures. Designed from the ground up, these accelerators often expose programming models that are distinct from GPUs. While hyperscalers and AI chip startups continue to innovate in this space, achieving broad operator coverage to support diverse models remains a major challenge. Additionally, an easy-to-use, high-level kernel programming language is important for rapid iteration of models and kernels. Triton, together with TorchInductor, addresses these issues on GPUs, but its viability on accelerators with different programming models has yet to be established. In this work, we present the first production-scale application of Triton on a custom ML accelerator, MTIA-2i, developed by Meta. To support MTIA-2i, we develop a new compiler backend that targets it, introduce enhancements to TorchInductor code generation, and propose minimal language extensions that expose MTIA-specific architectural features. We demonstrate that Triton-MTIA kernels achieve performance competitive with expert-tuned C++ implementations. Leveraging these development efficiency gains, we successfully deployed manually written and Inductor-generated Triton kernels in production across approximately 60 different model types, accounting for 50% of layers and 47% of non-GEMM execution time for these models. Our results provide compelling evidence that DSLs like Triton can bridge the programming model gaps between ML frameworks, kernels, and custom accelerators, enabling rapid innovation and efficient deployment at scale.

Comments12 pages, 12 figures, to be published in IEEE Micro

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑