AI 中文总结
针对非齐次随机图双样本检验在非整数 $L_r$ 范数下的样本复杂度缺口,提出基于 Hölder 插值组合两个统计量的检验方法,达到推测最优速率,并通过模拟验证。
AI 中文摘要
检验两个网络群体是否具有相同的边概率是网络推断中的一个基本问题。其难度取决于用于度量差异的范数。对于非齐次 Erdős–Rényi (IER) 模型,已知对于每个整数 $L_r$ 范数以及 $1\le r<2$ 的最优样本复杂度。然而,对于非整数 $r>2$,已知的上界和下界不匹配,且下界被推测为紧的。我们研究了在顶点对齐情况下的双样本检验中的这一差距。我们提出了一种检验方法,该方法在同一数据上运行两个已发表的统计量(阶数分别为 $2$ 和 $\lceil r\rceil$),若任一统计量拒绝则拒绝原假设。其阈值来自 Hölder 插值,使得两个统计量具有相同的样本成本。我们证明了该检验达到了推测的速率。结合先前的结果,这表明对于每个固定的 $r\ge 1$,极小极大样本复杂度为 $n^{\max\{4/r-1,\\,2/r\}}/\epsilon^2$ 阶,即使分离度随 $n$ 变化也是如此。在 $n$ 介于 32 到 256 之间的模拟中,在显著性水平 $0.05$ 下达到 80% 功效所需的图数量随 $n$ 的增长速率与理论一致。例如,对于 $r=2.5$,拟合指数为 $0.78$,而理论值为 $0.8$。有趣的是,两个统计量如插值论证所建议的那样分工:当只有少数边变化时,高阶统计量更有效;而当许多边变化时,$L_2$ 统计量更有效。
英文摘要
Testing whether two populations of networks share the same edge probabilities is a basic problem in network inference. How hard it is depends on the norm used to measure the difference. For the inhomogeneous Erdős--Rényi (IER) model, the optimal sample complexity is known for every integer $L_r$ norm and for $1\le r<2$. For non-integral $r>2$, however, the known upper and lower bounds do not match, and the lower bound was conjectured to be tight. We study this gap for two-sample testing on aligned vertices. We propose a test that runs two published statistics, of orders $2$ and $\lceil r\rceil$, on the same data and rejects if either one rejects. Its thresholds come from Hölder interpolation, so that both statistics have the same sample cost. We prove that this test attains the conjectured rate. Combined with earlier results, this shows that for every fixed $r\ge1$ the minimax sample complexity is of order $n^{\max\{4/r-1,\,2/r\}}/ε^2$, even when the separation changes with $n$. In simulations with $n$ between 32 and 256, the number of graphs needed for 80\% power at level $0.05$ grows with $n$ at a rate consistent with the theory. For $r=2.5$, for example, the fitted exponent is $0.78$, against the theoretical value $0.8$. Interestingly, the two statistics split the work as the interpolation argument suggests: the higher-order statistic is more powerful when only a few edges change, and the $L_2$ statistic when many edges change.