arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ESR-HGNN:消除语义冗余以实现高效的小批量异构图神经网络推理

ESR-HGNN: Eliminating Semantic Redundancy for Efficient Mini-batch HGNN Inference

Dengke Han, Mingyu Yan, Duo Wang, Wenming Li, Xiaochun Ye, Dongrui Fan

arXiv 2608.17865首次发表:更新:

AI 中文总结

本研究提出ESR-HGNN,通过元路径前缀树和可复用性驱动的元路径分组消除语义冗余,实现采样性能提升一个数量级,大幅降低能耗并加速小批量推理。

AI 中文摘要

异构图神经网络(HGNNs)在处理异构图数据方面非常有效,已被广泛应用于关键领域。随着现实世界图数据规模不断扩大,对整个图执行直接推理变得越来越不可行,使得小批量方法成为标准方法。然而,在端到端HGNN推理中,基于元路径的小批量采样由于图结构不规则遍历引发的大量随机内存访问,构成了显著的性能瓶颈。现有采样范式因固有语义冗余导致过度冗余遍历,严重降低采样效率,进而导致小批量推理性能欠佳。本研究提出一种感知冗余的HGNN采样范式,利用元路径前缀树复用遍历路径,有效消除冗余内存访问;随后将其映射至名为ESR-HGNN的多通道硬件采样单元。此外,引入可复用性驱动的元路径分组技术,对元路径进行最优聚类,以最大化硬件通道内可复用的遍历路径,提升语义并行场景下的效率。大量实验结果表明,ESR-HGNN相较于CPU和GPU实现了平均一个数量级的采样性能提升,同时显著降低能耗;当与GPU及最先进的HGNN推理加速器集成时,还能在端到端小批量推理中实现大幅加速。

英文摘要

Heterogeneous graph neural networks (HGNNs) are highly effective in processing heterogeneous graph data and have been widely adopted in critical domains. As real-world graph data continues to scale, performing direct inference on entire graphs becomes increasingly infeasible, making mini-batch methods the standard approach. However, in end-to-end HGNN inference, metapath-based mini-batch sampling constitutes a significant performance bottleneck due to the extensive random memory accesses induced by the irregular traversal of graph structures. Existing sampling paradigms suffer from excessive redundant traversals caused by inherent semantic redundancy, severely degrading sampling efficiency and, consequently, leading to suboptimal mini-batch inference performance. In this work, we propose a redundancy-aware HGNN sampling paradigm that leverages a metapath trie to reuse traversal paths, effectively eliminating redundant memory accesses. We then map it onto a multi-channel hardware sampling unit denominated ESR-HGNN. Furthermore, we introduce a reusability-driven metapath grouping technique that optimally clusters metapaths to maximize reusable traversal paths within hardware channels, enhancing efficiency in scenarios with semantic parallelism. Extensive experimental results demonstrate that ESR-HGNN achieves an average sampling performance improvement of one order of magnitude over CPU and GPU, accompanied by significant energy savings. Additionally, it delivers substantial speedup in end-to-end mini-batch inference when integrated with GPU and state-of-the-art HGNN inference accelerator.

Comments14 pages, 12 figures, to apear in IEEE TPDS (just accepted)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑