发表机构
RIKEN iTHEMS; RIKEN Center for Advanced Intelligence Project (AIP); South China University of Technology; Columbia University(理化学研究所跨学科理论科学中心; 理化学研究所先进智能研究中心; 华南理工大学; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究部分观测稀疏图中未知采样率的估计问题,证明其等价于尾指数估计,并给出模块化归约方法,在13个网络上无需拟合即达21.7%误差。
AI 中文摘要
大型图通常只能部分获取:因预算而停止的爬取、面板数据或部分转储。当采样比例 $s$ 通过设计已知时,总边数可由 $\hat e=e_s/s^2$ 得出,无需模型。我们处理 $s$ 未知且总体规模已知的情况。我们的主要结果是一个归约:在稀疏可交换(graphex)模型下,期望非孤立比例满足 $n_s/n_1\to s^{1+\sigma}$,因此一旦尾指数 $\sigma$ 可估计,采样率即可估计,将其代回得到 $e_s(n_1/n_s)^{2/(1+\sigma)}$——相同的估计量,但设计量被推断出来。因此,在稀疏图中估计全局边基数,在期望意义上等价于尾指数估计,而二次图on估计量是 $\sigma=0$ 的情形:它因恒等式而非拟合失败(中位误差 $260\\%$ 对比 $27\\%$)。我们界定了代换的有限尺寸误差,并表明该归约对尾指数估计量是模块化的——使用已发表的闭式估计量,在 $13$ 个网络和 $39$ 个采样预算上无需任何拟合即可达到 $21.7\\%$ 的误差。拟合完整的 graphex 模型还能返回任意规模的度分布和生成对象,其中稀疏性是一个坐标,插值路径由模型决定而非人为选择。两个极限是精确的:秩一 graphex 的传递性由度分布固定,因此高聚类图不在该类中;在雪球或随机游走爬取下,这里所有方法均失效,基于设计的oracle最差(误差从 $7.8\\%$ 到 $588\\%$)。
英文摘要
A large graph is often available only in part: a crawl stopped by its budget, a panel, a partial dump. When the sampled fraction $s$ is known by design the total edge count follows from $\hat e=e_s/s^2$ and no model is needed. We treat the case where $s$ is unknown and the population size is known. Our main result is a reduction: under a sparse exchangeable (graphex) model the expected non-isolated fraction obeys $n_s/n_1\to s^{1+σ}$, so the sampling rate becomes estimable once the tail index $σ$ is, and substituting it back gives $e_s(n_1/n_s)^{2/(1+σ)}$ -- the same estimator, with the design quantity inferred. Estimating global edge cardinality in a sparse graph is therefore, in expectation, tail-index estimation, and the quadratic graphon estimator is the case $σ=0$: it fails by an identity rather than by a fit ($260\%$ median error against $27\%$). We bound the finite-size error of the substitution and show the reduction is \emph{modular} in the tail-index estimator --- filled with a published closed-form one it reaches $21.7\%$ over $13$ networks and $39$ sampling budgets with no fitting at all. Fitting a full graphex additionally returns the degree distribution at any size and a generative object, in a representation where sparsity is a coordinate and the interpolation path is dictated rather than chosen. Two limits are exact: rank-one graphexes have transitivity fixed by the degree profile, so high-clustering graphs lie outside the class; and under snowball or random-walk crawls every method here fails, the design-based oracle worst of all ($7.8\%$ to $588\%$).