HoTS:同质性感知的温度缩放用于图神经网络校准
HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration
浏览论文内容
中文总结 AI 辅助
针对图节点分类中现有校准方法忽略局部结构的问题,提出同质性感知温度缩放(HoTS),利用熵和局部同质性为每节点分配温度,在18个基准上取得最优ECE 4.79%。
中文摘要 AI 辅助
对于图节点分类,当置信度分数(通常是最大预测类别概率)用于对预测进行排序、将不确定节点转交人工审查或控制风险时,需要校准的类别概率。现有的后处理校准器要么应用单一的全局温度,要么使用缺乏原则性结构形式的图感知模块。我们研究局部图结构应如何进入节点级校准。我们的初步结果表明,当具有相同logits但不同局部同质性的节点需要不同的最优温度时,仅基于logits的温度规则是不够的。然后,我们分析了一个具有高斯特征和单层线性GCN的群体集中上下文随机块模型。在等距类均值下,类模板分数上的贝叶斯后验是一个温度缩放的softmax,其逆温度由同质性依赖的信号强度控制。在正信号同质性区域,所得温度近似与归一化局部同质性成反比。该定律启发了同质性感知的温度缩放(HoTS),这是一种简单的后处理校准器,它根据基于熵的logits集中度和估计的局部同质性为每个节点分配一个正标量温度。HoTS有三个温度参数,保留预测类别,并从校准数据中学习结构校正的强度。在18个节点分类基准、两个GNN骨干和八个校准基线上,HoTS实现了最佳的平均期望校准误差(ECE)4.79%,最佳平均排名,以及在选择性分类中最可靠的置信度排序。代码可在该https URL获取。
英文摘要
For graph node classification, calibrated class probabilities are needed when confidence scores, usually the maximum predicted class probability, are used to rank predictions, defer uncertain nodes to human review, or control risk. Existing post-hoc calibrators either apply one global temperature or use graph-aware modules without a principled structural form. We study how local graph structure should enter node-level calibration. Our first results show that a logit-only temperature rule is insufficient when nodes with identical logits but different local homophily require different optimal temperatures. We then analyze a population-concentration contextual stochastic block model with Gaussian features and a one-layer linear GCN. Under equidistant class means, the Bayes posterior over class-template scores is a temperature-scaled softmax whose inverse-temperature is governed by a homophily-dependent signal strength. In the positive-signal homophilic regime, the resulting temperature decreases approximately inversely with normalized local homophily. This law motivates Homophily-aware Temperature Scaling (HoTS), a simple post-hoc calibrator that assigns each node a positive scalar temperature from entropy-based logit concentration and estimated local homophily. HoTS has three temperature parameters, preserves the predicted class, and learns the strength of the structural correction from calibration data. Across 18 node-classification benchmarks, two GNN backbones, and eight calibration baselines, HoTS achieves the best mean Expected Calibration Error (ECE) of 4.79%, the best average rank, and the most reliable confidence ranking in selective classification. Code is available at https://github.com/inu0104/HoTS.