AI 中文总结
本文针对GPU内存效率滞后问题,介绍DGNA方法,通过微基准测试和数据分析揭示GPU内存层次结构的NUMA架构,包括测量延迟的方法、应用高斯混合模型,还将其应用于英伟达GPU,获得相关架构和策略细节,为该领域研究提供了新成果。
AI 中文摘要
图形处理单元(GPU)凭借其强大的并行处理能力,在游戏和人工智能等各个领域变得至关重要。随着GPU核心的显著进步,GPU内存效率却滞后了,导致瓶颈限制工作负载效率。为弥合这一差距,深入了解GPU内存架构,尤其是L2和DRAM内的非统一内存访问(NUMA)机制,对于优化应用、设计新架构和构建精确模拟器至关重要。然而,英伟达和AMD等供应商的最新GPU硬件仍是黑箱,研究人员难以了解其设计细节。本文介绍了DGNA,一种通过微基准测试和数据分析揭示GPU内存层次结构NUMA架构的方法。具体而言,我们提出一种不依赖架构固有指令测量L2缓存和DRAM延迟的方法,并应用高斯混合模型过滤异常值并准确确定延迟分布。我们将DGNA应用于英伟达的A100和H100 GPU,揭示了NUMA节点架构、SM-NUMA关系以及用于维护缓存一致性的NUMA感知内存分配策略。据我们所知,这是第一篇详细介绍GPU内存子系统内NUMA架构的论文。
英文摘要
Graphics Processing Units (GPUs), due to their immense parallel processing capabilities, have become essential across various fields, including gaming and artificial intelligence. With significant advancements in GPU cores, GPU memory efficiency has lagged, resulting in bottlenecks that can limit workload efficiency. To bridge this gap, a deep understanding of GPU memory architectures, particularly Non-Uniform Memory Access (NUMA) mechanisms within L2 and DRAM, is essential for optimizing applications, designing new architectures, and building accurate simulators. However, the latest GPU hardware from vendors like NVIDIA and AMD is still a black-box, making it challenging for researchers to understand the details of their design. In this paper, we introduce DGNA, a methodology designed to unveil the NUMA architecture of the GPU memory hierarchy through microbenchmarking and data analysis. Specifically, we propose an approach to measuring the latency of L2 caches and DRAM without relying on the intrinsic instructions of the architecture and apply a Gaussian mixture model to filter out outliers and accurately determine latency distributions. We apply DGNA on NVIDIA's A100 and H100 GPUs, revealing NUMA node architecture, SM-NUMA relationships, and NUMA-aware memory allocation strategies used to maintain cache coherence. To the best of our knowledge, this is the first paper to detail the NUMA architecture within the GPU memory subsystem.