关于离散GPU与融合GPU同时协同调度的初步研究
A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPU
浏览论文内容
中文总结 AI 辅助
本研究对比GH200超级芯片与H100 PCIe平台的CPU-GPU协同调度,以稀疏共轭梯度为案例,发现GH200可让更多工作划分具备竞争力,提升协同调度的性能与可编程性。
中文摘要 AI 辅助
CPU-GPU协同调度可让应用在两个处理单元上同时执行,但其效率取决于工作负载划分和内存架构。本初步研究在NVIDIA GH200超级芯片上评估协同调度,并与离散H100 PCIe平台对比。以稀疏共轭梯度(CG)为案例,评估三种内存管理范式(显式拷贝、统一内存、映射内存)下的多种工作划分方式。评估凸显了减少手动CPU-GPU数据移动带来的运行时间与可编程性权衡。结果显示,与H100 PCIe平台相比,GH200让多种混合CPU-GPU工作划分具备竞争力,还让统一内存适用于多种矩阵。这些结果表明,GH200这类集成CPU-GPU平台可提升协同调度工作负载的性能与可编程性。
英文摘要
CPU-GPU coscheduling enables simultaneous execution of an application across both processing units, but its efficiency depends on workload partitioning and memory architecture. This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform. Using sparse conjugate gradient (CG) as a case study, we assess various work divisions across three memory-management paradigms: explicit copy, managed memory, and mapped memory. Our evaluation highlights the run time and programmability tradeoffs of reducing manual CPU-GPU data movement. The results show that compared with the H100 PCIe platform, GH200 makes several hybrid CPU-GPU work divisions competitive and makes managed memory practical for several matrices. These results suggest that integrated CPU-GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.