arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

剖析芯片缩放如何打破GPU细粒度调度

Dissecting How Die Scaling Breaks GPU Fine-grained Scheduling

Xiaoze Fan, Jianhao Wang, Weihao Cui, Han Zhao, Zhuobin Huang, Yangjie Zhou, Yuxian Qiu, Shixuan Sun, Bingsheng He, Quan Chen, Minyi Guo

arXiv 2609.24270首次发表:更新:

AI 中文总结

本研究剖析芯片缩放导致的GPU物理不对称性,提出轻量级表征与不对称性感知调度方法,提升主流内核、多路复用LLM推理性能并减少性能波动。

AI 中文摘要

现代GPU在物理上不再对称。芯片缩放导致制造驱动的晶圆级筛选以及缓存和内存分区。前者产生芯片特定的计算拓扑,而后者导致非均匀内存访问。这些不对称性是显著的。忽视拓扑的计算单元分配可导致高达1.33倍的性能变化,而远程访问使HBM延迟增加高达67%,并使L2延迟几乎翻倍。然而,这些不对称性隐藏在GPU的逻辑资源抽象之后,并可能因芯片而异。我们开发了轻量级表征方法,以揭示每块芯片的计算拓扑和内存亲和性。然后,我们利用发现的信息使现有的细粒度调度具有不对称性感知,不仅考虑分配了多少资源,还考虑分配了哪些物理资源。在整GPU内核执行、应用内多路复用和应用间共置中,不对称性感知调度将主流内核性能提升高达1.22倍,多路复用LLM推理性能提升高达14.3%,并避免了高达1.33倍的性能变化。

英文摘要

Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning. The former creates chip-specific compute topologies, while the latter causes non-uniform memory access. These asymmetries are substantial. Topology-oblivious compute unit allocation can lead to up to 1.33x performance variation, while remote accesses increase HBM latency by up to 67% and nearly double L2 latency. However, these asymmetries are hidden behind the GPU's logical resource abstractions and can vary across chips. We develop lightweight characterization methods to uncover per-chip compute topology and memory affinity. We then use the discovered information to make existing fine-grained scheduling asymmetry-aware, considering not only how many resources are allocated but also which physical resources are assigned. Across full-GPU kernel execution, intra-application multiplexing, and inter-application co-location, asymmetry-aware scheduling improves mainstream kernels by up to 1.22x, multiplexed LLM inference by up to 14.3%, and avoids up to 1.33x performance variation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑