用于归纳任务的层状神经网络基准测试
Benchmarking Sheaf Neural Networks for Inductive Tasks
浏览论文内容
中文总结 AI 辅助
本文通过首个层状神经网络设计空间的系统基准测试,在14个归纳式数据集上开展1890项实验,明确了层状神经网络在归纳任务中的表现及关键设计选择,为其实际应用提供了指导。
中文摘要 AI 辅助
层状神经网络(SNNs)通过将标准图神经网络(GNNs)的标量边权重替换为节点茎之间可学习的、依赖边的限制映射,对消息传递进行了泛化。尽管SNNs具有强大的理论基础和良好的直推式结果,但它们几乎仅在直推式节点分类任务上进行评估,其在归纳式协议下的表现尚不明确。我们通过首个层状设计空间的系统基准测试来解决这一空白,在一种从不组装重型层状拉普拉斯矩阵的消息传递重构框架内,评估三种扩散机制(神经层状扩散、层状注意力、结合图注意力网络v2的层状注意力)、三种限制映射参数化方式、三种茎维度以及六种现代GNN架构组件,该框架可在跨图批处理下训练整个设计空间。在14个归纳式数据集上开展的1890项受控实验中,获得多项关键见解:限制映射是主导性设计选择,通用映射更优;更大的茎维度增加了容量但未提升长程连接能力;架构组件比整个层状特定设计空间更能解释性能差异。在匹配协议下,SNNs可迁移至归纳式设置,但未达到最强基线,差距取决于数据集;实际上,单一的层状配置可跨数据集泛化,因此应将精力更多用于调整周围的架构方案,而非层状算子本身。
英文摘要
Sheaf Neural Networks (SNNs) generalize message passing by replacing scalar edge weights of standard Graph Neural Networks (GNNs) with learnable, edge-dependent restriction maps between node stalks. Despite their strong theoretical foundations and promising transductive results, SNNs have been evaluated almost exclusively on transductive node classification, leaving their behaviour under inductive protocols unknown. We address this gap through the first systematic benchmark of the sheaf design space, evaluating three diffusion mechanisms (neural sheaf diffusion, sheaf attention, and sheaf attention with Graph Attention Network v2), three restriction-map parameterizations, three stalk dimensions, and six modern GNN architectural components, within a message-passing reformulation that never assembles the heavy sheaf Laplacian, making the full design space trainable under cross-graph batching. Across $1{,}890$ controlled experiments on 14 inductive datasets, multiple insights emerge: restriction maps are the dominant design choice and general maps are preferable, larger stalks add capacity but not long-range reach, architectural components explain more performance variation than the entire sheaf-specific design space itself. Under a matched protocol, SNNs transfer to inductive settings but do not reach the strongest baselines, with gaps being dataset-dependent. Practically, a single sheaf configuration can generalize across datasets, so effort is better spent tuning the surrounding architectural recipe than the sheaf operator itself.