发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该综述将AI加速器分为四类,分析其计算、内存及互连演进,并探讨通用性与专用性权衡在数据中心规模的设计挑战。
AI 中文摘要
快速增长的人工智能工作负载正推动对AI数据中心的巨额投资。本综述将工业AI加速器分为四种架构类别,并比较其计算和内存组织。它考察了节点级、机架级和机柜级互连如何支持集合通信,并追踪了加速器代际间的架构演进。分析将算术吞吐量的提升与精度、数据传递、执行协调、通信、供电和冷却的变化联系起来。它还讨论了由工作负载多样性、数据移动、基础设施约束和模型演进引起的未来设计挑战,展示了通用性与专用性之间的权衡如何从单个加速器扩展到数据中心级系统。GitHub: github.com/Yufeng98/AI-datacenter
英文摘要
Rapidly growing AI workloads are driving large investments in AI datacenters. This survey classifies industrial AI accelerators into four architectural categories and compares their compute and memory organizations. It examines how node-, rack-, and pod-scale interconnects support collective communication, and traces architectural evolution across accelerator generations. The analysis connects advances in arithmetic throughput with changes in precision, data delivery, execution coordination, communication, power delivery, and cooling. It also discusses future design challenges arising from workload diversity, data movement, infrastructure constraints, and model evolution, showing how the trade-off between generality and specialization extends from individual accelerators to datacenter-scale systems. GitHub: github.com/Yufeng98/AI-datacenter