发表机构
Microsoft Corporation(微软公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出Maia 200 AI加速器,基于软件定义本地访问数据流架构,在特定功耗与带宽下实现高FP4、FP8性能,可降低成本能耗、支持AI推理大规模并行,适用于下一代高性能计算系统。
AI 中文摘要
我们推出Maia 200,这是一款先进的AI加速器,在750W TDP和7TB/s HBM带宽下,可提供10145 Tflop/s的FP4性能和5072 Tflop/s的FP8性能。Maia代表了一类新型软件定义本地访问数据流架构(SDLA),该架构明确对数据流引擎进行编程,以协调高度专业化的存储器和数据移动引擎。这种方法将重点从当今以线程为中心的架构转向以数据移动为中心的架构,提高了效率和可扩展性。我们受Flynn分类法启发提出的数据管理分类法,凸显了SDLA如何应对现代AI计算中的挑战。Maia 200在支持AI推理工作负载的大规模并行性的同时,实现了显著的成本和能源节约,使其成为下一代高性能计算系统的极具吸引力的解决方案。
英文摘要
We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines. This approach shifts the focus from today's thread-centric to data-movement-centric architecture, improving efficiency and scalability. Our taxonomy of data management, inspired by Flynn's classification, highlights how SDLA addresses challenges in modern AI computing. Maia 200 achieves significant cost and energy savings while supporting massive parallelism for AI inference workloads, making it a compelling solution for next-generation high-performance computing systems.