arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06441cs.LGcs.DC

SNI-GNN:借助SmartNIC的全图GNN训练,采用网络内嵌入预测

SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction

Guofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu, Amelie Chi Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

SNI-GNN是借助SmartNIC的全图GNN训练系统,通过网络内嵌入预测减少通信,在保持精度的同时实现21%-45%通信量降低,1.3-3.6倍加速,可高效扩展至16个GPU。

中文摘要 AI 辅助

全图GNN训练精度高,但在多服务器集群上扩展性差,原因是节点间嵌入交换量大且不规则。本文提出SNI-GNN,一种借助SmartNIC的全图训练系统,通过在网络内预测远程嵌入减少通信量同时保持精度。SNI-GNN在SmartNIC上部署轻量线性趋势预测器,用于优化缓存的历史嵌入,结合基于重要性的边界节点采样策略,以及带中间结果复用的异步DPU-GPU数据流水线。我们给出误差和收敛界,证明在有界二阶动态下预测器偏差可被控制,且在不精确梯度下实现标准非凸收敛。在NVIDIA BlueField-3上实现的SNI-GNN与最先进的全图系统集成,通信量减少21%至45%,相较于BNS-GCN实现1.3至3.6倍的端到端加速,相较于基线SANCUS实现最高1.29倍加速,精度损失≤0.01,在边数达数千万的图上可高效扩展至16个GPU。这些结果表明,基于SmartNIC的网络内预测是分区和压缩技术的实用补充,适用于大规模通信高效的全图GNN训练。

英文摘要

Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We present SNI-GNN, a SmartNIC-assisted full-graph training system that reduces communication while preserving accuracy by predicting remote embeddings in-network. SNI-GNN deploys a lightweight linear-trend predictor on SmartNICs to refine cached historical embeddings, coupled with an importance-based boundary-node sampling policy and an asynchronous DPU--GPU data pipeline with intermediate-result reuse. We provide error and convergence bounds showing that predictor bias remains controlled under bounded second-order dynamics and yields standard non-convex convergence with inexact gradients. Implemented on NVIDIA BlueField-3, SNI-GNN integrates with state-of-the-art full-graph systems, cuts communication by 21--45\%, achieves 1.3--3.6$\times$ end-to-end speedups over BNS-GCN and up to 1.29$\times$ over baseline SANCUS, with accuracy loss $\leq 0.01$, and scales efficiently to 16 GPUs on graphs with up to tens of millions of edges. These results indicate SmartNIC-based in-network prediction is a practical complement to partitioning and compression techniques for communication-efficient full-graph GNN training at scale.

补充信息

↑