发表机构
College of Sciences, China Jiliang University(中国计量大学理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出融合节点中心性与局部相似性指标的链接预测框架,利用PageRank和DomiRank构建DR-MD系列,在九个数据集上显著优于基线,DR-RA平均AUC达0.7084。
AI 中文摘要
局部相似性指标因其计算成本低而被广泛用于复杂网络中的链接预测;然而,在稀疏网络中,它们为所有缺乏共同邻居的节点对分配零分数,这严重限制了其预测能力。一种自然的补救措施是将节点中心性指标与局部相似性指标相融合:前者提供节点对的全局重要性,而后者捕捉细粒度的局部拓扑,两者可以在统一框架内组合成互补的分数。本文使用PageRank和DomiRank作为两种代表性的中心性度量,构建了一个中心性-局部相似性融合框架。首先将Charikhi提出的基于PageRank的融合推广到七种经典的局部相似性指标,并在九个真实世界网络数据集上系统验证了其改进的普适性。此外,引入DomiRank中心性,在统一加权系数下构建DR-MD系列融合指标,克服了基于PageRank的融合需要逐指标调整权重的缺点。五折交叉验证以及Wilcoxon符号秩检验的结果表明,在统一实验协议下,所有DR-MD指标在所有九个数据集上均一致优于相应的局部基线和其PR-MD对应物(p=0.002),并且改进在近临界参数平台内对σ和加权系数的扰动保持稳健;特别是,DR-RA实现了0.7084的平均AUC,超过了Katz和RWR等全局方法以及几种先进的相似性指标。该框架本质上是可扩展的,其融合范式可以直接推广以将其他节点中心性指标与局部相似性指标耦合。
英文摘要
Local similarity indices assign zero scores to node pairs without common neighbors, which limits link prediction in sparse networks; fusing node centrality with local similarity is a common remedy, but existing fusion studies use heterogeneous protocols and the robustness of their gains is unclear. Within a unified piecewise fusion framework (multiplicative modulation when local information is sufficient, small-dose completion when it is absent), we show that the dynamic range of the centrality product governs the modulation mechanism: the PageRank product is of order O(n^-2), so its factor degenerates to a near-identity map. Under a fair protocol with exhaustive negative-sample comparison on eight real-world networks, PageRank fusion therefore yields no consistent significant gain for six of the seven local indices; its single exception, a gain of about +0.035 on the near-tree-like wiki-Vote network, comes entirely from completion, indicating that the completion gain depends on both the fraction of zero-score node pairs and the standalone predictive power of the centrality product. The min-max normalized DomiRank product, with an O(1) range, instead gives stable gains under unified parameters (modulation weight 5, completion coefficient 0.1): all seven fused indices improve significantly at the fold level (p <= 3.3e-3), and a 10x5 repeated cross-validation confirms robustness to the randomness of fold partitions. DR-RA reaches an average AUC of 0.9257, surpassing Katz, LNB, CN2D, CNC, CND, SimRank, CCPA, Gravity and CNPop, and comparable to RWR. The design rules obtained for the two mechanisms (completion coefficient at most 0.3, modulation weight in [0.5,12], competition intensity near critical) provide a reproducible protocol benchmark and quantitative parameter design principles for the centrality x local-similarity fusion paradigm.
Comments29 pages, 6 figures, 11 tables, 44 references. v3: adds scale-treatment control experiments (Section 4.4), a degree-corrected benchmark (Section 5.7) and a data/code availability statement, with the corresponding abstract, limitations and conclusion statements; corrects numeric ranges and one per-dataset claim