arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

聚合关系数据中的社区检测:可辨识性与谱恢复

Community Detection from Aggregated Relational Data: Identifiability and Spectral Recovery

Andrew Davison, Owen G. Ward

arXiv 2610.03937首次发表:更新:

发表机构

Simon Fraser University(西蒙菲莎大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究聚合关系数据下度校正随机块模型的社区检测,给出可辨识性条件,证明谱聚类在特定条件下精确恢复社区,并引入二阶查询理论。

AI 中文摘要

聚合关系数据将完全观测的网络替换为压缩版本,其中调查受访者被问及诸如“你的朋友中有多少属于G组?”的问题。因此,我们仅观察到邻接矩阵A与固定设计O的乘积矩阵R = OA。在本文中,我们研究了在度校正随机块模型下,这种压缩对社区检测的影响。首先,我们给出了在何种条件下能够恢复完整网络参数的保证,以及在基于E[R]的对角校正版本条件下仅恢复社区的保证,这些条件可用于指导实验设计。其次,我们证明当存在K个平衡社区且满足nρ_n ≳ K^2 + K log n时,在“均匀分布”或近正交设计下,使用单链路的谱聚类能够精确恢复社区,并更一般地刻画了设计如何影响收敛速率。我们还引入了二阶“朋友的朋友”查询,并发展了相应的理论。

英文摘要

Aggregated relational data replaces a fully observed network with a compressed version, where survey respondents are asked questions such as ``How many of your friends belong to group $G$?''. We consequently only observe a matrix $R = OA$ of the adjacency matrix $A$ with a fixed design $O$. In this paper we understand the effects of this compression on community detection under degree-corrected stochastic block models. First, we give guarantees for when it is possible to recover the full network parameters, and then only just the communities under conditions on a diagonal corrected version of $\mathbb{E}[R]$, which can be used to inform experimental design. Secondly, we show that spectral clustering with single linkage exactly recovers communities when $nρ_n\gtrsim K^2+K\log n$ when there are $K$ balanced communities under ``uniformly spread'' or near-orthogonal designs, and characterize more generally how the design influences rates of convergence. We also introduce second-order ``friends-of-friends'' queries and develop the corresponding theory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑