arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25273cs.LG

HeAD-CP:用于图神经网络的异质性感知扩散共形预测集

HeAD-CP: Heterophily-Aware Diffused Conformal Prediction Sets for Graph Neural Networks

Phan Binh Nguyen Lam, Nguyen Thai Anh

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对图神经网络共形预测中扩散方法的不足,提出HeAD-CP,通过GNN softmax导出的无标签局部同质性估计确定系数,有三个变体,在多个基准测试中表现优于DAPS,有效解决异质性问题。

中文摘要 AI 辅助

共形预测(CP)提供无分布的不确定性量化,其向图的扩展是一个活跃的研究方向。扩散自适应预测集(DAPS)是一种广泛使用的图感知扩散基线,沿边以均匀系数λ传播自适应预测集(APS)不一致分数。我们发现这种设计的一个基本缺点:均匀低通扩散预先假定图的同质性,在异质图上证明是有害的,相对于普通APS,平均预测集大小增加高达10.6%。为了缓解这一问题,我们提出了HeAD-CP,这是一族节点级扩散变体,其系数由从GNN softmax导出的无标签局部同质性估计确定。三个变体,即带符号γ、边兼容性和带校正的DAPS基线,分别在极端异质性、中等异质性和中高同质性下最有效,并且都保留了边际覆盖保证。在十个基准上,HeAD-CP家族在每个数据集上都保持在普通APS或以下水平,但DAPS在六个数据集上超过了APS。该家族的事后预言机在8/10个数据集上在p<0.01(配对Wilcoxon)时比DAPS有所改进,在异质图上增益最大(在Texas上为10.3%);在DAPS仍然获胜的两个同质数据集(CiteSeer、PubMed)上,它最多保留0.002的边际优势,在CiteSeer上无统计学意义(p=0.23)。设计一个接近该预言机的校准无标签选择器是主要的突出实证问题。

英文摘要

Conformal prediction (CP) provides distribution-free uncertainty quantification, and its extension to graphs is an active research direction. Diffused Adaptive Prediction Sets (DAPS) is a widely used graph-aware diffusion baseline, propagating Adaptive Prediction Sets (APS) non-conformity scores along edges with a uniform coefficient $λ$. We identify a fundamental shortcoming of this design: the uniform low-pass diffusion presupposes graph homophily and proves detrimental on heterophilic graphs, enlarging the mean prediction-set size by up to 10.6% relative to plain APS. To mitigate this, we propose HeAD-CP, a family of node-wise diffusion variants whose coefficients are determined by a label-free local-homophily estimate derived from the GNN softmax. Three variants, namely signed-$γ$, edge-compatibility, and a DAPS-baseline-with-correction, are most effective at extreme heterophily, intermediate heterophily, and moderate-to-high homophily, respectively, and all preserve the marginal coverage guarantee. On ten benchmarks, the HeAD-CP family stays at or below plain APS on every dataset, while DAPS exceeds APS on six. The post-hoc oracle over the family improves over DAPS on 8/10 datasets at $p<0.01$ (paired Wilcoxon), with the largest gains on heterophilic graphs (10.3% on Texas); on the two homophilic datasets where DAPS still wins (CiteSeer, PubMed), it retains a marginal advantage of at most 0.002, statistically insignificant on CiteSeer ($p=0.23$). Designing a calibrated label-free selector that approaches this oracle is the main outstanding empirical question.

补充信息

↑