发表机构
Center for Complex Biological Systems; University of California, Irvine(复杂生物系统中心; 加州大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计贝叶斯图对齐中得分收敛与边际准确性的差距,比较多种诊断方法,发现赋值敏感审计有效但受预算和参考一致性限制。
AI 中文摘要
贝叶斯图对齐估计对应概率,但对齐得分轨迹的收敛并不一定意味着对应边际分布的准确性。我们在来自四个源家族的240对新的精确图、240对包含20至100个顶点的较大图对,以及一个独立的60案例精确实现检查中,审计了这一差距。在显式边翻转似然下,我们比较了三种采样器以及基于得分、边际、指示器、类别和分类器的诊断方法。对于精确信息采样器,边际不一致性在错误判别上优于得分R-hat,但对于普通局部采样,其改进效果不确定。基于赋值的R*和短指示器面板具有竞争力;没有一种诊断在所有采样器和端点中占主导地位。在较大规模下,诊断预测的是后续边际变化而非后验误差,且分类性能取决于漂移阈值。不重叠窗口和保留链检查减弱但保留了正相关性。240个原始参考集中仅有22个通过一致性筛选。在四十个失败选择的案例中,八倍SMC粒子升级未能解决不一致性,而额外的再生化有所帮助。更长的信息运行仍不稳定。一个基本的可行对齐界限表明,在集中的100顶点案例中,SMC和信息链得分严重缺乏代表性,这与近似参考共识无关。我们还展示了共同起始链在精确边际误差接近0.967时仍具有接近零的不一致性。这些结果支持赋值敏感的审计,同时指出了有限预算、诊断排名和参考一致性作为准确性证据的局限性。
英文摘要
Bayesian graph alignment estimates correspondence probabilities, but convergence of an alignment-score trace need not imply accurate correspondence marginals. We audit this gap on 240 new exact graph pairs from four source families, 240 larger pairs with 20-100 vertices, and a separate 60-case exact implementation check. Under an explicit edge-flip likelihood, we compare three samplers and score, marginal, indicator, categorical, and classifier-based diagnostics. Marginal disagreement improves error discrimination over score R-hat for the exact informed sampler, but its improvement for vanilla local sampling is uncertain. Assignment-based R* and short indicator panels are competitive; no diagnostic dominates across samplers and endpoints. At larger sizes, diagnostics predict subsequent marginal changes, not posterior error, and classification performance depends on the drift threshold. Disjoint-window and held-out-chain checks attenuate but preserve positive associations. Only 22 of 240 original reference sets pass an agreement screen. On forty failure-selected cases, eightfold SMC particle escalation does not resolve disagreement, whereas additional rejuvenation helps. Longer informed runs remain unstable. An elementary feasible-alignment bound demonstrates severely unrepresentative SMC and informed-chain scores in concentrated 100-vertex cases, independently of approximate reference consensus. We also exhibit common-start chains with near-zero disagreement despite exact marginal error near .967. These results support assignment-sensitive auditing while identifying limits of finite budgets, diagnostic rankings, and reference agreement as evidence of accuracy.
Comments19 pages, 4 figures. Codes: https://github.com/Mirsohi/Graph-Alignment