AI 中文总结
本研究剖析芯片缩放导致的GPU物理不对称性,提出轻量级表征与不对称性感知调度方法,提升主流内核、多路复用LLM推理性能并减少性能波动。
AI 中文摘要
现代GPU在物理上不再对称。芯片缩放导致制造驱动的晶圆级筛选以及缓存和内存分区。前者产生芯片特定的计算拓扑,而后者导致非均匀内存访问。这些不对称性是显著的。忽视拓扑的计算单元分配可导致高达1.33倍的性能变化,而远程访问使HBM延迟增加高达67%,并使L2延迟几乎翻倍。然而,这些不对称性隐藏在GPU的逻辑资源抽象之后,并可能因芯片而异。我们开发了轻量级表征方法,以揭示每块芯片的计算拓扑和内存亲和性。然后,我们利用发现的信息使现有的细粒度调度具有不对称性感知,不仅考虑分配了多少资源,还考虑分配了哪些物理资源。在整GPU内核执行、应用内多路复用和应用间共置中,不对称性感知调度将主流内核性能提升高达1.22倍,多路复用LLM推理性能提升高达14.3%,并避免了高达1.33倍的性能变化。
英文摘要
Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning. The former creates chip-specific compute topologies, while the latter causes non-uniform memory access. These asymmetries are substantial. Topology-oblivious compute unit allocation can lead to up to 1.33x performance variation, while remote accesses increase HBM latency by up to 67% and nearly double L2 latency. However, these asymmetries are hidden behind the GPU's logical resource abstractions and can vary across chips. We develop lightweight characterization methods to uncover per-chip compute topology and memory affinity. We then use the discovered information to make existing fine-grained scheduling asymmetry-aware, considering not only how many resources are allocated but also which physical resources are assigned. Across full-GPU kernel execution, intra-application multiplexing, and inter-application co-location, asymmetry-aware scheduling improves mainstream kernels by up to 1.22x, multiplexed LLM inference by up to 14.3%, and avoids up to 1.33x performance variation.