arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无顶点对应的随机图双样本检验

Two-Sample Testing for Random Graphs without Vertex Correspondence

Soham Dan

arXiv 2610.07503首次发表:更新:

AI 中文总结

针对无顶点对应的随机图双样本检验,研究了所需样本量及统计量能力,发现不对齐代价为$t^{-2}$,带符号三角形计数最优,并给出精确有效检验。

AI 中文摘要

两个图群体经常需要在顶点之间没有任何对应关系的情况下进行比较,例如当网络来自不同社区,或当图生成模型针对留出图进行评估时。我们研究了这种不对齐的双样本检验需要多少图,以及哪些图统计量可以检测哪些差异。对于Erdős--Rényi零假设和一种保持每个期望度不变的植入双块差异,我们证明当每图信噪比为$t<1$时,每组需要且仅需要$m\asymp t^{-3}$个图。带符号三角形计数达到此速率,且下界对每个图大小都成立。当顶点对齐时,$m\asymp t^{-1}$个图就足够,因此不对齐造成的代价约为$t^{-2}$倍。当三角形信号抵消时,速率变为$t^{-4}$,需要4-环。由树构建的统计量在两种假设下具有完全相同的期望,基于有限个此类统计量的检验渐近地没有功效。在图极限中,此类包括度分布和消息传递图神经网络特征。对于非恒定零假设,一般差异在一阶可见,简单的模体检验达到对齐的样本量阶,表明不对齐主要对低阶不可见的差异代价高昂。我们还给出了每组一个或两个图的精确有效检验,但以功效为代价。在我们的模拟中,拟合指数接近预测值,在带符号三角形需要约$65$个图的设置中,基于度和随机GNN的评估指标保持其水平。

英文摘要

Two populations of graphs often have to be compared without any correspondence between their vertices, for instance when networks come from different communities, or when a graph generative model is evaluated against held-out graphs. We study how many graphs such an unaligned two-sample test needs, and which graph statistics can detect which differences. For an Erdős--Rényi null and a planted two-block difference that leaves every expected degree unchanged, we show that $m\asymp t^{-3}$ graphs per group are necessary and sufficient when the per-graph signal-to-noise ratio is $t<1$. Signed triangle counts attain this rate, and the lower bound holds for every graph size. With aligned vertices $m\asymp t^{-1}$ graphs suffice, so misalignment costs a factor of order $t^{-2}$. When the triangle signal cancels, the rate becomes $t^{-4}$ and $4$-cycles are needed. Statistics built from trees have exactly the same expectation under both hypotheses, and tests based on finitely many of them have asymptotically no power. In the graphon limit, this class includes degree distributions and message-passing graph neural network features. For a non-constant null, a generic difference is visible at first order, and a simple motif test attains the aligned order of sample size, suggesting that misalignment is costly mainly for differences that are invisible at low orders. We also give an exactly valid test for one or two graphs per group, at a cost in power. In our simulations, the fitted exponents are close to the predicted ones, and degree-based and random-GNN evaluation metrics stay at their level in a setting where signed triangles need about $65$ graphs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑