arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

越过内存墙,进入指令墙:GPU数据处理中的新瓶颈

Over the Memory Wall, Into the Instruction Wall: The New Bottleneck in GPU Data Processing

Sven Hepkema, Bowen Wu, Christos Kozyrakis, Yannis Chronis, Gustavo Alonso

arXiv 2608.13696首次发表:更新:

AI 中文总结

该研究针对GPU数据处理中内存带宽提升未使cuDF内核成比例加速的问题,构建Valk工具分析发现新瓶颈为指令墙,提出提升缓存利用等三项优化建议以发挥GPU潜力。

AI 中文摘要

数据中心GPU随着新一代高带宽内存(HBM)的采用,内存带宽提升了一个数量级。与此同时,GPU数据库系统日益受到关注,许多系统基于开源GPU关系算子库cuDF构建。此前,查询性能受限于内存带宽,但内存带宽的提升并未带来cuDF内核的成比例加速。为探究性能未能跟上的原因,我们构建了性能分析工具Valk,它结合了多个分析器的数据。我们在两种硬件性能极端的GPU(L4和GH200)上对运行TPC-H内存数据库的cuDF进行分析。GH200的内存带宽是L4的13.4倍,指令吞吐量是L4的2.5倍,但运行TPC-H仅快5.2倍。分析表明,内存带宽提升后,内核变为计算受限。基于分析,我们提出三项建议,以在内存墙消失时充分发挥GPU在关系型工作负载上的潜力:内核需1)更高效地利用缓存;2)提升占用率和/或指令级并行度;3)减少每次内存访问的指令执行量。

英文摘要

Datacenter GPUs have seen an order-of-magnitude increase in memory bandwidth with the adoption of newer generations of HBM. Meanwhile, GPU database systems are gaining traction, many building on cuDF, an open-source library of GPU relational operators. Previously, query performance was bound by memory bandwidth, but the increase in memory bandwidth has not resulted in a proportional speedup of cuDF kernels. To investigate why performance has not kept up, we built Valk, a performance analysis tool that combines data from multiple profilers. We profile cuDF running TPC-H in-memory on two extremes of hardware capability, the L4 and GH200 GPUs. The GH200 has 13.4$\times$ the memory bandwidth and 2.5$\times$ the instruction throughput of the L4, yet is only 5.2$\times$ faster in running TPC-H. Our analysis shows that when memory bandwidth is increased, kernels become compute bound. From our analysis, we make three recommendations to fully utilize the GPUs' potential for relational workloads when the memory wall is removed: kernels need to 1) make more efficient use of caches, and 2) increase occupancy and/or instruction level parallelism, and 3) execute fewer instructions per memory access.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑