FlashGPU-sim:为现代架构和AI工作负载实现GPU建模
FlashGPU-sim: Enabling GPU Modeling for Modern Architectures and AI Workloads
- Chinese University of Hong Kong(香港中文大学)
- Zhejiang University(浙江大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
FlashGPU-sim是一个开源、周期精确的GPU模拟器,支持现代架构和AI工作负载,通过建模异步数据移动、张量核心等特性,实现5.24%的MAPE和7.86倍多线程加速。
AI中文摘要:
随着人工智能日益普及,现代AI系统由紧密的软硬件协同设计循环所塑造。较新的GPU暴露了诸如异步数据移动、张量核心流水线和细粒度同步等特性,高性能内核会积极利用这些特性,而新兴应用行为也日益影响下一代硬件设计。不幸的是,最新的NVIDIA GPU开源模拟器专注于大约六年前的架构和软件栈。因此,它们无法支持现代编译器栈(例如Triton)生成的许多最先进的AI内核,也无法准确建模这些内核所依赖的硬件特性。结果,架构师缺乏一个可信的平台来分析这一飞轮中的瓶颈或评估未来AI系统的设计权衡。为了弥合这一差距,我们提出了FlashGPU-sim,一个面向现代AI工作负载的开源、执行驱动、周期精确的GPU模拟器。FlashGPU-sim忠实地建模了诸如异步数据移动、细粒度同步、张量核心执行和分布式共享内存等现代硬件特性。一个Triton提取前端允许直接模拟优化的AI算子而无需手动移植,而多线程执行使大规模软硬件协同设计变得实用。在RTX 5090、H100和B200上的131个工作负载配置中,FlashGPU-sim实现了5.24%的周期级MAPE,而多线程模拟在16个主机线程下达到了7.86倍的加速。一个H100案例研究进一步证明了其在微架构设计探索中的实用性。
英文摘要:
As AI becomes increasingly ubiquitous, modern AI systems are shaped by a tight software-hardware co-design loop. Later GPUs expose features such as asynchronous data movement, tensor core pipelines, and fine-grained synchronization that high-performance kernels aggressively exploit, while emerging application behaviors increasingly influence the next generation of hardware design. Unfortunately, the latest open-source simulators for NVIDIA GPUs focus on architectures and software stacks from roughly six years ago. Therefore, they cannot support many state-of-the-art AI kernels generated by modern compiler stacks, e.g. Triton, or accurately model the hardware features they depend on. As a result, architects lack a credible platform for analyzing bottlenecks in this flywheel or evaluating design trade-offs for future AI systems. To bridge this gap, we present FlashGPU-sim, an open-source, execution-driven, cycle-accurate GPU simulator for modern AI workloads. FlashGPU-sim faithfully models modern hardware features such as asynchronous data movement, fine-grained synchronization, tensor-core execution, and distributed shared memory. A Triton extraction front-end allows direct simulation of optimized AI operators without manual porting, while multi-threaded execution makes large-scale software-hardware co-design practical. Across 131 workload configurations on RTX 5090, H100, and B200, FlashGPU-sim achieves a cycle-level MAPE of 5.24%, while multi-threaded simulation reaches a 7.86x speedup with 16 host threads. An H100 case study further demonstrates its utility for microarchitectural design exploration.