arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Maia 200:用于大规模AI加速的软件定义数据流系统

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu, Li Zhang, Torsten Hoefler

arXiv 2608.24664首次发表:更新:

发表机构

Microsoft Corporation(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出Maia 200 AI加速器,基于软件定义本地访问数据流架构,在特定功耗与带宽下实现高FP4、FP8性能,可降低成本能耗、支持AI推理大规模并行,适用于下一代高性能计算系统。

AI 中文摘要

我们推出Maia 200,这是一款先进的AI加速器,在750W TDP和7TB/s HBM带宽下,可提供10145 Tflop/s的FP4性能和5072 Tflop/s的FP8性能。Maia代表了一类新型软件定义本地访问数据流架构(SDLA),该架构明确对数据流引擎进行编程,以协调高度专业化的存储器和数据移动引擎。这种方法将重点从当今以线程为中心的架构转向以数据移动为中心的架构,提高了效率和可扩展性。我们受Flynn分类法启发提出的数据管理分类法,凸显了SDLA如何应对现代AI计算中的挑战。Maia 200在支持AI推理工作负载的大规模并行性的同时,实现了显著的成本和能源节约,使其成为下一代高性能计算系统的极具吸引力的解决方案。

英文摘要

We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines. This approach shifts the focus from today's thread-centric to data-movement-centric architecture, improving efficiency and scalability. Our taxonomy of data management, inspired by Flynn's classification, highlights how SDLA addresses challenges in modern AI computing. Maia 200 achieves significant cost and energy savings while supporting massive parallelism for AI inference workloads, making it a compelling solution for next-generation high-performance computing systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑