arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HARP-ME:基于GPU的闭包驱动精确诱导子图枚举

HARP-ME: Closure-Driven Exact Induced Motif Enumeration on GPUs

Ashwina Kumar, Rupesh Nasre

arXiv 2607.12074首次发表:更新:

AI 中文总结

研究针对GPU上精确诱导子图枚举的挑战,提出HARP-ME框架,通过闭包感知编译、诱导签名重用等技术,优化锚点选择,实现高效枚举,在多个数据集上速度远超其他方法,证明其在GPU子图枚举中的优势。

AI 中文摘要

精确诱导子图枚举是图挖掘中的基本操作,在GPU上具有挑战性。本文提出HARP-ME框架用于精确枚举连通诱导四节点子图。它引入闭包感知编译,通过考虑遍历成本等因素选择锚点。还引入诱导签名重用技术,利用候选前沿信息和紧凑约束识别可重用完成状态。对于超出GPU内存的图,通过规范锚点所有者规则确保跨分区精确计数。实验表明HARP-ME是最快的,缓存命中率在64%到76%之间,主机到设备的数据传输开销更低。

英文摘要

Exact induced motif enumeration is a fundamental operation in graph mining, but it remains challenging on GPUs because candidate expansion is irregular, repeated set intersections dominate execution, and induced counting must consider both the presence and absence of edges. We present HARP-ME, which stands for Hierarchical Anchor-Reuse Partitioned Motif Enumeration. It is a GPU framework for the exact enumeration of connected induced four-node motifs. HARP-ME introduces closure-aware compilation, which selects a set of anchors for explicit enumeration by considering traversal cost, algebraic derivation benefits, expected state reuse, and the additional halo overhead caused by graph partitioning. HARP-ME also introduces induced-signature reuse. This technique identifies reusable completion states using candidate-frontier information together with compact constraints representing both adjacency and non-adjacency relationships. For graphs that exceed GPU memory, a canonical anchor-owner rule ensures exact counting across partitions with overlapping halo regions. We do not claim that the individual graphlet closure identities are new. Instead, our contribution is their integration into a GPU execution framework that reduces both explicit candidate expansion and repeated induced-edge checks. Across six social, web, biological, and synthetic graphs, HARP-ME is the fastest among the evaluated methods. It achieves up to 2.11 times speedup over Pangolin, up to 1.83 times speedup over partitioned PBE, and up to 10.73 times speedup over the evaluated CPU baseline. Detailed measurements show cache-hit rates between 64% and 76%, along with substantially lower host-to-device data transfer overhead than PBE-style partitioning. These results demonstrate that optimizing anchor selection for algebraic derivation benefits can complement traversal-based and reuse-based GPU motif enumeration techniques.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑