发表机构
Capital Normal University; Institute of Science Tokyo; Beijing Institute of Technology(首都师范大学; 东京科学大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对GCN部署的GPU工作负载不平衡问题,提出双度量划分与自适应核执行的DualGCN框架,在12个真实数据集上较现有工具实现2.13-3.8倍的平均加速。
AI 中文摘要
图卷积网络(GCN)广泛应用于社交、引用、电商网络等大型图结构数据处理,但它们的部署受限于不规则内存访问和严重的GPU工作负载不平衡问题。这些挑战来自两个维度:一是幂律度分布导致的宽度不平衡,二是异构邻域导致的深度不平衡。本文提出DualGCN,一种通过双度量图划分和自适应核执行解决这两个维度问题的GPU加速框架。DualGCN结合反映聚合宽度的节点度,以及通过匿名随机游走估计的、捕捉多跳连通性和访问深度的邻域密度。这种混合工作负载度量可将大图划分为稀疏和密集区域,同时将工作负载不平衡从线性复杂度降低到对数复杂度。随后,DualGCN为不同分区选择特定执行策略:稀疏分区采用 warp 级并行和合并内存访问,而密集分区则利用指令级并行隐藏延迟、提升GPU利用率。在12个真实世界图数据集上的实验表明,DualGCN可持续加速GCN计算,相较于cuSPARSE、GNNAdvisor和ACCEL,分别实现2.53倍、3.8倍和2.13倍的平均加速比。这些结果证明,联合优化图划分与核执行是处理大规模图和社交网络工作负载的有效方案。
英文摘要
Graph Convolutional Networks (GCNs) are widely used for large graph-structured data, including social, citation, and e-commerce networks, but their deployment is constrained by irregular memory access and severe GPU workload imbalance. These challenges arise in two dimensions: width imbalance from power-law degree distributions and depth imbalance from heterogeneous neighborhood connectivity.We present DualGCN, a GPU acceleration framework addressing both dimensions through dual-metric graph partitioning and adaptive kernel execution. DualGCN combines node degree, reflecting aggregation width, with neighborhood density estimated by anonymous random walks, capturing multihop connectivity and access depth. This hybrid workload metric enables connectivity-aware partitioning of large graphs into sparse and dense regions while reducing workload imbalance from linear to logarithmic complexity. DualGCN then selects partition-specific execution strategies: sparse partitions use warp-level parallelism and coalesced memory access, whereas dense partitions exploit instruction-level parallelism to hide latency and improve GPU utilization. Experiments on twelve real-world graph datasets show that DualGCN consistently accelerates GCN computation, achieving average speedups of 2.53x, 3.8x, and 2.13x over cuSPARSE, GNNAdvisor, and ACCEL, respectively. These results demonstrate that jointly optimizing graph partitioning and kernel execution provides an effective solution for processing large-scale graph and socialnetwork workloads.
CommentsThis research work has already been accepted by WISE'2026