AI 中文总结
针对隐藏独立级联模型下的逆问题,提出基于模拟的摊销估计器SAGE-HC,利用门控图神经网络从症状观测恢复节点级传播参数,在环状图上优于DMP基准。
AI 中文摘要
信息和传染病通过社交网络传播,但驱动这些传播的概率在缺乏每个节点的激活时间的情况下难以估计。应用场景很少提供这些数据,因此推断必须通过终端感染状态的间接且带噪声的代理进行。我们在固定且已知的图上,在隐藏独立级联(HIC)模型下研究这一逆问题,每个节点有一个传播概率,而非单一的全局速率,因此未知数的数量随节点数增长。种子集合和观测参数已知,而激活时间和终端感染状态是潜在的,观测数据的似然需要对每个传播结果进行边缘化。我们提出了一种基于模拟的摊销估计器,无需重建单个潜在级联即可恢复完整的节点级参数向量。重复的种子条件症状观测被总结为症状感知级联特征(SACF),其结合了经验症状统计与邻域和结构信息。SACF通过SAGE-HC映射到参数,SAGE-HC是一种置换等变门控图神经网络,其学习到的门控衰减由假阳性和假阴性污染的邻域消息,在模拟HIC实现上的训练产生了一个可复用的逆映射。作为相同隐藏观测下的基准,我们将动态消息传递学习框架扩展到HIC发射模型。在合成和实证图上,两种方法按拓扑分离。DMP在树上高度准确,而SAGE-HC在具有噪声终端症状的异构、环状图上表现显著更好。
英文摘要
Information and infectious diseases spread through social networks, but the spreading probabilities driving them are hard to estimate without per-node activation times. Applications seldom supply these, and inference must instead proceed through indirect and noisy proxies for the terminal infection states. We study this inverse problem on a fixed, known graph under the Hidden Independent Cascade (HIC) model, with one spreading probability per node rather than a single global rate, so the number of unknowns scales with the number of nodes. Seed sets and observation parameters are known, while activation times and terminal infection states are latent, and the observed-data likelihood requires marginalizing over every spreading outcome. We propose a simulation-based amortized estimator that recovers the full node-level parameter vector without reconstructing individual latent cascades. Repeated seed-conditioned symptom observations are summarized as Symptom-Aware Cascade Features (SACF), which combine empirical symptom statistics with neighborhood and structural information. SACF are mapped to parameters by SAGE-HC, a permutation-equivariant gated graph neural network whose learned gates attenuate neighborhood messages corrupted by false positives and false negatives, and training on simulated HIC realizations yields a reusable inverse map. As a benchmark under the same hidden observations, we extend the Dynamic Message Passing learning framework to the HIC emission model. On synthetic and empirical graphs the two methods separate by topology. DMP is highly accurate on trees, whereas SAGE-HC is substantially better on heterogeneous, loopy graphs under noisy terminal symptoms.