LGNNIC:利用SmartNIC加速大规模GNN训练
LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs
- Technion Israel Institute of Technology(以色列理工学院)
- Hewlett Packard Enterprise Labs(慧与科技实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出LGNNIC架构,利用SmartNIC卸载分布式GNN训练的预处理任务,经PoC系统验证,邻居采样和量化技术可显著降低通信开销,实现训练加速。
AI中文摘要:
图神经网络(GNN)已广泛应用于自然科学、社交网络分析、芯片设计、推荐系统等领域。然而,随着图规模的增长,将其全部存储并在单节点CPU-GPU系统上处理变得愈发不切实际。一种有前景的方法是将图分布在多个远程内存节点上,但这会引入一个主要瓶颈:训练期间的节点间网络拥塞。为解决该问题,我们提出LGNNIC,一种新颖的节点间系统架构,利用与远程内存节点共存的SmartNIC(现代系统中已有的配置)来减少分布式GNN训练中的通信开销。LGNNIC将关键预处理任务卸载到SmartNIC,减少传输到计算(训练)节点的数据量,缓解网络拥塞。我们在预处理阶段的SmartNIC上引入两种互补技术:执行小批量采样的邻居采样(Neighbor Sampling),以及对采样批次的量化(Quantization)。为在不同通信基础设施下评估LGNNIC,我们设计了优化的低开销基于DMA的同步机制,以及作为基准的高开销基于套接字(Socket)的替代方案。我们使用由1个搭载NVIDIA BlueField-2 SmartNIC的远程内存节点和1个搭载A100 GPU的计算节点组成的概念验证(PoC)系统,在标准GNN工作负载和采样超参数下评估核心SmartNIC卸载机制。远程节点上的邻居采样和量化在大多数配置中均实现了显著的训练加速:邻居采样在套接字和DOCA-DMA下分别实现最高62.4倍和17.5倍的加速,主要源于数据传输时间减少;量化通过减少数据传输,分别提供最高3.6倍和1.3倍的额外加速。
英文摘要:
Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems. However, as graph sizes grow, storing and processing them entirely on a single-node CPU-GPU system becomes increasingly impractical. A promising approach is to distribute the graph across multiple remote memory nodes, though this introduces a major bottleneck: inter-node network congestion during training. To address this, we propose LGNNIC, a novel inter-node system architecture that leverages SmartNICs co-located with remote memory nodes-a configuration already available in modern systems-to reduce communication overhead in distributed GNN training. LGNNIC offloads key preprocessing tasks to SmartNICs, reducing the volume of data transferred to computational (training) nodes and alleviating network congestion. We introduce two complementary techniques executed on the SmartNICs during the preprocessing phase: Neighbor Sampling, which performs mini-batch sampling, and Quantization of the sampled batches. To evaluate LGNNIC under different communication infrastructures, we designed both an optimized low-overhead DMA-based synchronization mechanism and a high-overhead socket-based alternative used as a benchmark. We evaluate the core SmartNIC offloading mechanisms across standard GNN workloads and sampling hyperparameters using a proof-of-concept (PoC) system comprising one remote-memory node with an NVIDIA BlueField-2 SmartNIC and one compute node with an A100 GPU. Both Neighbor Sampling and Quantization on the remote node demonstrated substantial training speedups in most configurations. Neighbor Sampling achieved up to 62.4x and 17.5x speedups with Sockets and DOCA-DMA, respectively, primarily due to reduced data transaction time. Quantization provided additional speedups of up to 3.6x and 1.3x, respectively, by reducing data transfer.