arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于公共邻居计数的不确定图上代表性可能世界发现

Common-Neighbor-Count-Based Representative Possible World Finding on Uncertain Graphs

Chengjie Gu, Xiaoliang Xu, Yuxiang Wang, Kai Yao, Mengzhao Wang, Tianxing Wu, Yingjie Xia, Xiangyu Ke

arXiv 2607.23085首次发表:更新:

AI 中文总结

研究不确定图上基于公共邻居计数的代表性可能世界发现问题,将其从保留节点级统计扩展到保留成对结构关系,证明问题NP难,开发两阶段算法并加速细化、设计终止方法,实验验证算法在多种挖掘任务尤其是公共邻居相关任务中的有效性。

AI 中文摘要

代表性可能世界(RPW)是从不确定图$\mathcal{G}$派生的确定性图,其中指定结构特征在$\mathcal{G}$中接近其期望值。作为$\mathcal{G}$的代理,RPW可让传统确定性算法直接在其上执行针对该特征的挖掘任务,避免对$\mathcal{G}$进行计算昂贵的枚举或采样。现有研究主要关注单个节点特征,而许多挖掘任务依赖两节点间公共邻居数量这一两对特征。为此研究基于公共邻居计数的代表性可能世界(CRPW)问题,将RPW从保留节点级统计扩展到保留成对结构关系。证明该问题是NP难的,开发两阶段基本算法,用高效整数计数策略加速细化,设计基于Beta的自适应终止方法。在真实不确定图上的大量实验证明了算法在各种挖掘任务上的有效性,尤其在与公共邻居相关的任务中性能最佳。

英文摘要

A representative possible world (RPW) is a deterministic graph derived from an uncertain graph $\mathcal{G}$ where a designated structural feature closely approximates its expected value in $\mathcal{G}$. Serving as a proxy for $\mathcal{G}$, the RPW allows conventional deterministic algorithms to be directly executed on it for mining tasks targeting this feature, thereby avoiding computationally expensive enumeration or sampling on $\mathcal{G}$. Existing studies on RPWs primarily focus on individual node features, e.g., degree or triangle degree. However, many mining tasks, such as link prediction, critically rely on the number of common neighbors between two nodes, which is a pairwise feature. To bridge this gap, we study the \underline{C}ommon-neighbor-count-based \underline{R}epresentative \underline{P}ossible \underline{W}orld (CRPW) problem, extending RPWs from preserving node-level statistics to preserving pairwise structural relationships. The problem seeks the possible world that best preserves the expected numbers of common neighbors between node pair, and we prove that is NP-hard. To address it, we develop a two-stage basic algorithm that quickly initializes a possible world and then refines it iteratively. We next accelerate the refinement by replacing its costly floating-point evaluation with an efficient integer counting strategy, as the refinement only requires determining whether a change is beneficial, rather than computing its exact magnitude. Moreover, we design a Beta-based adaptive termination method to automatically stop the refinement once the desired quality of the possible world is reached, preventing over- or under-execution. Extensive experiments on real-world uncertain graphs demonstrate the effectiveness of our algorithms on diverse mining tasks. Especially on common-neighbor-related tasks, we achieve the best performance among all compared methods.

CommentsFull version; 16 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑