发表机构
University of Cambridge; Ruhr University Bochum; Microsoft(剑桥大学; 波鸿鲁尔大学; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长程图基准缺乏可验证长程性的问题,提出四个公理及TRIP/GRIP构造框架,可生成可证明的长程任务并提供测试误差下界,用于审计现有基准并揭示过压缩的拓扑与计算瓶颈。
AI 中文摘要
关于图神经网络(GNN)中过压缩与长程交互之间联系的实证主张,只有在用于验证这些主张的基准测试真正需要长程交互时才能被信任。事实上的标准——长程图基准(Long Range Graph Benchmark)——已被反复证明会被调优后的短程模型所饱和,而现有的合成替代方案则与特定拓扑结构绑定。因此,在任意图上缺乏有原则的长程性证书。这种状况反映了缺乏对长程基准的精确刻画。我们通过引入四个可验证的公理来解决这一根本性差距:可预测性(Predictability)、紧致性(Tightness)、严格k-范围(Strictly $k$-Range)和拓扑不变性(Topology-Invariance),任何声称测试k跳交互的任务都必须满足这些公理。我们正式证明,违反其中任何一条都会导致破坏从任务中得出的结论的失败模式。基于这些公理,我们引入了TRIP(真正范围交互问题)及其推广GRIP(一般范围交互问题),这些构造性过程通过从稳定分布中提取特征,将任何图转化为可证明的长程任务。此外,通过构造,GRIP具有闭式、逐范围的最大似然预言机,该预言机提供了在任何基准上可用的首个先验逐范围测试误差下界。使用我们的框架,我们:(i)审计了4个常见的长程基准,并识别出它们相对于我们公理的失败模式;(ii)在TRIP实例化的拓扑上,我们发现一种流行的曲率概念与GNN性能不相关,支持拓扑瓶颈与计算瓶颈的区分;(iii)我们展示了一个新基准的过压缩度量所衡量的因素超出了纯粹的长程性。使用该框架和复现实验的代码已发布在https URL。
英文摘要
Empirical claims about the connection between over-squashing and long-range interactions in GNNs, can only be trusted if the benchmarks used to validate them genuinely require long-range interactions. The de-facto standard, the Long Range Graph Benchmark, has been repeatedly shown to be saturated by tuned short-range models, with existing synthetic alternatives being tied to specific topologies. As such, there is a lack of principled certificate of long-rangedness on arbitrary graphs. This state reflects the absence of a precise characterization of long-ranged benchmarks. We address this fundamental gap by introducing four verifiable axioms: Predictability, Tightness, Strictly $k$-Range, and Topology-Invariance, that any task claiming to test $k$-hop interactions must satisfy. We formally prove that violating any one of them admits failure modes that undermine conclusions drawn from the task. Based on these axioms, we introduce TRIP (Truly Ranged Interactions Problem) and its generalisation GRIP (Generally Ranged Interactions Problem), constructive procedures that turn any graph into a provably long-ranged task by drawing features from stable distributions. Moreover, by construction, GRIP admits a closed-form, per-range Maximum-Likelihood oracle that yields the first a priori per-range lower bound on test error available on any benchmark. Using our framework, we: (i) audit 4 common long-range benchmarks and identify their failures modes with respect to our axioms; (ii) on TRIP-instantiated topologies, we find a popular notion of curvature is uncorrelated with GNN performance, supporting topological-vs-computational bottleneck distinction; and (iii) we show that a novel benchmark's over-squashing measures factors beyond pure long-rangedness. Code to use the framework and reproduce experiments is released https://github.com/ferranhernandezc/graph-grip.
CommentsPublished at the Conference on Neural Information Processing Systems (NeurIPS 2026). Track on Evaluations and Datasets