发表机构
University of California, Los Angeles; Block; Mila – Quebec AI Institute(加州大学洛杉矶分校; Block公司; Mila魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出EB-GAD,一种免训练图异常检测框架,将正常性建模为图感知GOU松弛,通过经验贝叶斯拟合先验,将评分转化为有限时域控制能量,在11个基准上无标签取得9个最佳AUROC。
AI 中文摘要
节点级图异常检测(GAD)识别其属性与交互偏离主导图规律性的节点。现有GAD模型通过架构、消息传递、重构或对比目标以及调整后的评分族,间接编码正常性和异常评分。这纠缠了图信任度(图结构应在多大程度上定义正常性)、图谱加权和异常评分选择,导致评分在不同图机制下成本高昂、不透明且不稳定。我们提出EB-GAD(经验贝叶斯GAD),一个免训练框架,将正常性建模为面向图滤波模板的图感知广义Ornstein-Uhlenbeck(GOU)松弛。经验贝叶斯从残差场似然中拟合图精度;随后GOU将评分转化为闭式有限时域控制能量,即沿图谱松弛将特征中性节点引导至其观测终点所需的最小努力。扫描松弛时域和终点容差产生一组共享一个拟合先验的评分:平衡Mahalanobis评分是一个极限,而有限时域控制能量和尺度归一化比率评分揭示了静态平衡评分可能掩盖的异常。一个无标签选择器根据特征同质性、边密度和特征维度选择评分族,然后根据拟合零偏差和秩稳定性对候选进行排序。在11个基准上,且在每一步无标签的情况下,EB-GAD在9个基准上达到最佳或并列最佳AUROC:四个金融欺诈网络(最大达370万个节点)、YelpChi和Amazon评论图、Weibo、Reddit和Facebook,优势最高达21.7个百分点。在BlogCatalog和ACM上排名第二。
英文摘要
Node-level graph anomaly detection (GAD) identifies nodes whose attributes and interactions deviate from dominant graph regularities. Existing GAD models encode normality and anomaly scoring indirectly through architectures, message passing, reconstruction or contrastive objectives, and tuned score families. This entangles graph trust (how strongly graph structure should define normality), graph-spectral weighting, and anomaly-score choice, yielding scores that are costly, opaque, and unstable across graph regimes. We propose EB-GAD (Empirical-Bayes GAD), a training-free framework that models normality as graph-aware generalized Ornstein-Uhlenbeck (GOU) relaxation toward a graph-filtered template. Empirical Bayes fits the graph precision from the residual-field likelihood; the GOU then turns scoring into a closed-form finite-horizon control energy, the minimum effort to steer a feature-neutral node to its observed endpoint along graph-spectral relaxation. Sweeping relaxation horizon and endpoint tolerance yields a bank of scores that share one fitted prior: equilibrium Mahalanobis scoring is one limit, while finite-horizon control-energy and scale-normalized ratio scores reveal anomalies that static equilibrium scoring can mask. A label-free selector chooses the score family from feature homophily, edge density, and feature dimension, then ranks candidates by fitted-null deviation and rank stability. On 11 benchmarks and without labels at any step, EB-GAD has the best or tied-best AUROC on 9: the four financial fraud networks (up to 3.7M nodes), the YelpChi and Amazon review graphs, Weibo, Reddit and Facebook, with margins of up to 21.7 points. It is second on BlogCatalog and ACM.
CommentsPaper already accepted at Neurips