GPU-Tile-Sim:用于大语言模型软硬件协同设计的以瓦片为中心的GPU仿真框架
GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型中GPU内核给性能模型带来的挑战,提出GPU-Tile-Sim框架,以瓦片图表示内核执行,设计前后端,在多种工作负载上评估,精度高,还扩展到Blackwell验证其对设计分析的有效性。
中文摘要 AI 辅助
现代大语言模型工作负载越来越依赖通过软硬件协同设计优化的GPU内核。这些内核通过细粒度依赖调度和计算与内存重叠实现高性能,给现有GPU性能模型带来新挑战。指令驱动模拟器适应架构演变成本高,分析模型又过于粗略。我们提出GPU-Tile-Sim,一种用于大语言模型软硬件协同设计的以瓦片为中心的GPU仿真框架。关键在于现代大语言模型内核性能受依赖结构而非单个指令延迟影响更大。GTSim将内核执行表示为一个 warp 级瓦片图,设计了自动瓦片图前端和图驱动仿真后端。在代表性的GEMM、注意力和端到端大语言模型推理工作负载上评估GTSim,在A100和H100上,GTSim实现了高性能建模精度(平均绝对百分比误差,MAPE,1.22% - 8.71%)。还将GTSim扩展到Blackwell并初步验证,证明其在分析软件和架构设计选择方面的有效性。
英文摘要
Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grained dependency scheduling and computation-memory overlap. As such, they incur new challenges on existing GPU performance models. Instruction-driven simulators are costly to adapt to evolving architectures, while analytical models are too coarse to capture kernels' characteristics. We propose GPU-Tile-Sim, a tile-centric GPU simulation framework for LLM hardware-software co-design. The key insight is that modern LLM kernel performance is governed less by individual instruction latency than by the dependency structure that controls execution order and overlap. Accordingly, GTSim represents kernel execution as a warp-level tile graph whose nodes capture tile-level operations and whose edges encode data and ordering constraints. Using this representation, we design an automatic tile-graph frontend and a graph-driven simulation backend. We evaluate GTSim on representative GEMM, attention, and end-to-end LLM inference workloads. On A100 and H100 across both conventional and highly optimized kernels, GTSim achieves high performance-modeling accuracy (MAPE, Mean Absolute Percentage Error, 1.22%--8.71%). We further extend GTSim to Blackwell with preliminary validation, and demonstrate its effectiveness in analyzing software and architectural design choices.