arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13655cs.LGstat.MEstat.ML

分布偏移下归纳图的在线贝叶斯节点分类

Online Bayesian Node Classification on Inductive Graphs under Distribution Shift

Jinwen Xu, Gonzalo Mateos Buckstein, Qin Lu

首次发表
浏览论文内容

中文总结 AI 辅助

针对演化图分布偏移下的节点分类,提出在线变分贝叶斯最后一层(GVBLL)方法,联合训练编码器与近似后验,测试时在线更新,在五个基准上均取得最佳准确率与负对数似然。

中文摘要 AI 辅助

在演化图上,节点分类器必须满足两个关键要求:在分布偏移下对新到达节点进行归纳泛化,以及针对安全敏感应用提供校准的不确定性。标准图神经网络(GNN)通常只训练一次,且无法满足这两个要求。我们通过将随机最后一层参数置于确定性GNN编码器之上,改编了贝叶斯最后一层(BLL)模型,用于不确定性量化。分类所需的分类softmax似然破坏了高斯共轭性,因此训练后验和测试时流式更新均无闭式解。为解决这两个挑战,我们引入变分贝叶斯最后一层(VBLL)目标,通过最大化带有蒙特卡洛期望对数似然的证据下界,联合训练编码器和近似最后一层后验。在测试时,我们冻结编码器,并对最后一层后验应用在线拉普拉斯更新。该更新对应于具有指数遗忘和Kullback-Leibler锚定到训练后验的幂先验贝叶斯模型。在分布偏移下的五个节点分类基准上,在线GVBLL是唯一在每个数据集上均达到最佳准确率和负对数似然的方法。与最强的非GVBLL基线相比,它在Cora上准确率提高多达17个百分点,在ogbn-arxiv上提高14个百分点,同时在校准方面与MC Dropout、Deep Ensembles、Temperature Scaling和高斯过程分类器保持竞争力。

英文摘要

On evolving graphs, node classifiers must satisfy two key requirements: inductive generalization to newly arriving nodes under distribution shift and calibrated uncertainty for safety-sensitive applications. Standard graph neural networks (GNNs) are typically trained once and address neither requirement. We adapt the Bayesian last-layer (BLL) model by placing random last-layer parameters on top of a deterministic GNN encoder for uncertainty quantification. The categorical softmax likelihood required for classification breaks Gaussian conjugacy, so neither the training posterior nor the test-time streaming update has a closed-form solution. To address both challenges, we introduce a variational Bayesian last-layer (VBLL) objective that jointly trains the encoder and an approximate last-layer posterior by maximizing an evidence lower bound with a Monte Carlo expected log-likelihood. At test time, we freeze the encoder and apply an online Laplace update to the last-layer posterior. This update corresponds to a power-prior Bayesian model with exponential forgetting and a Kullback-Leibler anchor to the training posterior. Across five node-classification benchmarks under distribution shift, online GVBLL is the only method to achieve the best accuracy and negative log-likelihood on every dataset. It improves accuracy by up to 17 percentage points on Cora and 14 percentage points on ogbn-arxiv over the strongest non-GVBLL baseline, while remaining competitive in calibration with MC Dropout, Deep Ensembles, Temperature Scaling, and Gaussian-process classifiers.

发表机构

  • University of Georgia(佐治亚大学)
  • University of Rochester(罗切斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑