面向AI时代的下一代异步分布式GPU架构设计
Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era
浏览论文内容
中文总结 AI 辅助
针对AI时代GPU架构模拟基础设施滞后的问题,提出可建模Ampere等现代GPU的周期级模拟框架,经H100验证精度达标,用于评估小芯片拓扑等新兴设计轨迹。
中文摘要 AI 辅助
机器学习工作负载的快速发展从根本上改变了GPU硬件,推动架构向多芯片模块(MCM)拓扑、异步执行原语以及持久的多阶段内核行为演进。尽管发生了这些变化,周期级模拟基础设施却滞后,缺乏原生能力来建模现代GPU的物理非均匀性以及最先进AI工作负载的大规模规模。为了弥合这一差距,我们提出了一个周期级模拟框架,旨在准确建模包括Ampere、Hopper和Blackwell在内的现代GPU代。该模拟器经过物理硅的严格验证,在H100 GPU上实现了99%的皮尔逊相关系数和13.4%的平均绝对周期误差。利用该基础设施,我们开展架构案例研究,评估新兴设计轨迹,包括小芯片拓扑扩展、扩展SRAM容量与带宽以及GPU间预取策略。
英文摘要
The rapid evolution of machine learning workloads has fundamentally transformed GPU hardware, driving architectures toward Multi-Chip Module (MCM) topologies, asynchronous execution primitives, and persistent, multi-phase kernel behaviors. Despite these shifts, cycle-level simulation infrastructure has lagged behind, lacking the native capability to model the physical non-uniformity of modern GPUs alongside the massive scale of state-of-the-art AI workloads. To bridge this gap, we present a cycle-level simulation framework designed to accurately model modern GPU generations, including Ampere, Hopper, and Blackwell. Rigorously validated against physical silicon, the simulator achieves a 99% Pearson correlation coefficient and a 13.4% mean absolute cycle error on the H100 GPU. Utilizing this infrastructure, we conduct architectural case studies to evaluate emerging design trajectories, including chiplet topology scaling, expanded SRAM capacity and bandwidth, and inter-GPU prefetching strategies.