AI 中文总结
该研究针对多芯片GPU扩展系统中未被充分研究的非统一网络访问(NUNA)问题,提出感知NUNA的路由(NAR)与布局(NAP)方法,可将机器学习推理的集体通信加速最高达1.8倍,每输出token时间最大提升28%。
AI 中文摘要
图形处理单元(GPU)架构规模不断增长,以满足日益提升的计算与内存需求。随着GPU规模增大,插槽内的线路传输延迟显著增加。尽管过往研究已针对单个插槽内的计算与内存局部性进行优化,但GPU间通信的空间影响尚未得到充分研究。我们引入术语非统一网络访问(Non-Uniform Network Access,NUNA),用以描述多GPU系统中这一新兴优化维度,尤其聚焦于机器学习推理中常见的延迟敏感型集体通信。首先,我们强调需采用感知NUNA的路由(NUNA-aware Routing,NAR),在大规模扩展网络拓扑中选择经优化的、感知空间的GPU间路径;其次,我们引入感知NUNA的布局(NUNA-aware Placement,NAP),将线程块与数据放置在I/O附近以优化GPU间流量。我们证明,仅NAP优化即可实现比未感知局部性的基线高最多1.5倍的集体通信加速;将NAP与NAR结合后,相比未感知局部性的基线,集体通信加速可达最多1.8倍,这使得机器学习推理中每输出token的平均时间提升7%(最大提升28%)。
英文摘要
Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly. While previous research has optimized for compute and memory locality within a socket, the spatial impact on inter-GPU communication has not been well-studied. We introduce the term non-uniform network access (NUNA) to describe this emerging optimization dimension in multi-GPU systems. We specifically focus on latency-sensitive collective communication, common in machine learning inference. First, we highlight the need for NUNA-aware routing (NAR), choosing optimized, spatially-aware inter-GPU paths in large scale-up network topologies. Second, we introduce NUNA-aware placement (NAP), placing threadblocks and data near I/O to optimize the inter-GPU traffic. We demonstrate that the NAP optimizations alone offer up to 1.5x collective speedups over a locality-unaware baseline. Combining NAP with NAR yields up to 1.8x faster collectives over the locality-unaware baseline. This leads to 7% mean (28% max) time per output token speedup in machine learning inference.
Comments15 pages total, 11 pages body, 15 figures, 2 tables, 1 algorithm