arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于图神经网络可扩展训练的两级域分解AdaGrad方法

Two-level domain-decomposition AdaGrad method for scalable training of graph neural networks

Laurynas Varnas, Julien Herrmann, Alexander Heinlein, Serge Gratton, Alena Kopaničáková

arXiv 2608.22575首次发表:更新:

发表机构

Institut de Recherche en Informatique de Toulouse; Artificial and Natural Intelligence Toulouse Institute(图卢兹计算机科学研究院; 图卢兹人工智能与自然智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对图神经网络分布式训练的高成本问题,提出两级域分解AdaGrad方法,经实验可降本4-8倍且性能提升最高22%。

AI 中文摘要

图神经网络(GNNs)已成为从图结构数据中学习的强大框架,但其高效训练仍具挑战性,尤其在分布式计算环境中。这一挑战源于消息传递的使用,其将所有图节点耦合,导致优化步骤昂贵、内存需求高且通信开销大。为缓解这些局限,我们提出了AG2m的新型域分解(DD)变体,AG2m是一种增强了二阶曲率信息和动量的AdaGrad方法,记为DD-AG2m。所提出的DD-AG2m在原始(全局)图上的AG2m优化与分区图上的AG2m优化之间交替进行。为以更低成本融入全局信息,我们进一步引入两级变体(2DD-AG2m),该方法对通过在每个子域内随机采样节点获得的粗图执行全局优化步骤。涵盖图分类、节点级回归和时空预测任务的数值实验表明,所提出的DD方法在达到相同预测性能时,可将计算成本降低4至8倍。此外,在固定计算成本下,与基线AG2m相比,它们可将GNNs的预测性能提升高达22%。

英文摘要

Graph neural networks (GNNs) have emerged as a powerful framework for learning from graph-structured data. However, their efficient training remains challenging, particularly in distributed computing environments. This challenge arises from the use of message passing, which couples all graph nodes, leading to expensive optimization steps, high memory requirements, and substantial communication overhead. To alleviate these limitations, we propose a novel domain-decomposition (DD) variant of AG2m, an AdaGrad method enhanced with second-order curvature information and momentum, denoted by DD-AG2m. The proposed DD-AG2m alternates between AG2m optimization on the original (global) graph and AG2m optimization on the partitioned graphs. To incorporate global information at reduced cost, we further introduce a two-level variant (2DD-AG2m) that performs global optimization steps on a coarse graph obtained by randomly subsampling nodes within each subdomain. Numerical experiments spanning graph classification, node-level regression, and spatiotemporal forecasting tasks demonstrate that the proposed DD methods reduce the computational cost required to achieve the same predictive performance by a factor of 4-8. Moreover, for the fixed computational cost, they improve the predictive performance of GNNs by up to 22% compared with the baseline AG2m.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑