发表机构
Beihang University; Qingdao Research Institute, Beihang University; State Key Laboratory of Complex and Critical Software Environment, Beihang University(北京航空航天大学; 北京航空航天大学青岛研究院; 北京航空航天大学复杂关键软件环境国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MegaGraph提出首个面向图Transformer的自动化混合并行框架,通过图感知上下文、异构流水线和混合数据并行及成本模型搜索,解决大规模训练的内存和负载问题,实现高达77.8%内存降低和4.51倍加速。
AI 中文摘要
图Transformer(GT)通过克服传统图神经网络(GNN)的深度限制和过平滑问题,提供了优越的表示能力。然而,将GT扩展到大规模图会引发关键瓶颈。具体而言,注意力分数矩阵及其相关的拓扑感知偏置矩阵共同导致显著的每层内存开销,而繁重的图嵌入层导致严重的工作负载不平衡。这些特性是GT训练所特有的,并且针对传统GNN或Transformer设计的并行技术无法解决这些问题,因此需要专门的解决方案。本文介绍了MegaGraph,这是第一个专为高效GT训练设计的自动化混合并行框架。MegaGraph设计了三种专门策略,即图感知上下文并行、异构流水线并行和混合数据并行,以支持在大规模图上的高效训练。然而,协调这三种并行策略会产生指数级的配置空间。为了解决这一复杂性,自动搜索引擎通过Profile-Model-Search工作流利用精确的成本模型来识别最优并行配置。评估表明,MegaGraph能够在大规模图上进行训练,而最先进的基线方法因内存不足(OOM)错误而失败。该框架将每设备峰值内存减少高达77.8%,并实现高达4.51倍的训练加速,同时保持模型准确性。
英文摘要
Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated topology-aware bias matrix jointly incur significant per-layer memory overhead, and heavy graph embedding layers result in severe workload imbalances. These characteristics are unique to GT training and are not addressed by parallelism techniques designed for either conventional GNNs or Transformers, making a dedicated solution necessary. This paper introduces MegaGraph, the first automated hybrid parallelism framework designed for efficient GT training. MegaGraph designs three specialized strategies, namely graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism, to support efficient training on large-scale graphs. However, coordinating these three parallelism strategies yields an exponentially large configuration space. To address this complexity, an automatic search engine leverages precise cost models via a Profile - Model - Search workflow to identify the optimal parallelism configuration. Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors. The framework reduces per-device peak memory by up to 77.8\% and achieves up to 4.51$\times$ training speedup while maintaining model accuracy.
Comments8 pages,10 figures, accepted by ICCD2026