arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sobek:流形等变张量积卷积

Sobek: Streaming Equivariant Tensor Product Convolutions

Vladimir Chorošajev, Cédric Bény

arXiv 2607.18074首次发表:更新:

发表机构

Cortex Discovery(皮质发现公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究等变图神经网络中传统张量积卷积实现的问题,提出通过重新关联操作优化的流形公式,在Sobek中实现后经评估,相比传统方法有显著速度提升、内存减少,能处理更大工作负载。

AI 中文摘要

等变图神经网络在图边反复应用边条件张量积卷积。传统实现会生成特定于边的权重、消息和伴随量,导致张量积工作空间和内存流量随图大小和算子宽度迅速增长。这限制了可行的工作负载,阻碍更大问题充分利用GPU。我们表明这些边大小的中间量是执行调度的产物,而非等变算子的要求。通过重新关联径向投影、球谐耦合和图聚合,边局部积可直接被消耗到有界接收端状态。由此产生的流形公式保留了全连接多重性混合,并贯穿前向、反向和双反向。我们在Sobek(一个生成的CUDA后端)中实现了此公式,并在边缩放机制和各种特征结构上进行评估。在所有75次容量匹配比较中,Sobek在两个算子族和所有三个微分阶数上都更快,加速比从1.2倍到49.7倍不等,并将峰值分配内存减少多达99%。它还能执行比OpenEquivariance容量高出两个数量级的工作负载,同时保持接近峰值的吞吐量。这些结果表明,边缩放张量积工作空间是传统调度的属性,而非等变卷积本身的属性。

英文摘要

Equivariant graph neural networks repeatedly apply edge-conditioned tensor-product convolutions over graph edges. Conventional implementations materialize edge-specific weights, messages, and adjoints, causing tensor-product workspace and memory traffic to grow rapidly with graph size and operator width. This limits feasible workloads and can prevent larger problems from fully utilizing the GPU. We show that these edge-sized intermediates are artifacts of the execution schedule, not requirements of the equivariant operator. By reassociating radial projection, spherical-harmonic coupling, and graph aggregation, edge-local products can be consumed directly into bounded receiver-side state. The resulting streaming formulation preserves fully connected multiplicity mixing and extends through forward, backward, and double backward. We implement this formulation in Sobek, a generated-CUDA backend, and evaluate it across edge-scaling regimes and varied feature structures. Across two operator families and all three differentiation orders, Sobek is faster in all 75 capacity-matched comparisons, with speedups ranging from $1.2\times$ to $49.7\times$, and reduces peak allocated memory by up to 99\%. It also executes workloads up to two orders of magnitude beyond OpenEquivariance's capacity while retaining near-peak throughput. These results show that edge-scaled tensor-product workspace is a property of the conventional schedule, not of equivariant convolution itself.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑