发表机构
Ulsan National Institute of Science and Technology; Soonchunhyang University(蔚山科学技术院; 顺天乡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出表示感知模块度RAM,通过距离校准零模型解决经典模块度中零模型不匹配问题,实验证明其优于多种基线方法。
AI 中文摘要
模块度是社区检测的标准目标函数,它将观测到的社区内部连接性与零模型期望进行比较。现代节点嵌入捕获了超越原始拓扑的结构和语义邻近性,但在保留经典基于度的零模型的同时,用表示诱导的亲和性替换邻接矩阵,使得观测项依赖于几何结构,而期望项与几何无关。这造成了零模型不匹配:目标函数可能奖励那些已经由表示邻近性预期的亲和性。本文提出了RAM,一种具有距离校准零模型的表示感知模块度目标函数。RAM通过同时校准度异质性和表示距离来保持观测减期望的原则。我们分析了其与经典模块度的关系,并开发了分裂式和多层次贪心优化器。在七个真实世界数据集和合成可扩展性测试上的实验表明,RAM优于仅基于拓扑的方法、直接嵌入聚类方法和基于深度学习的基线方法,在经典模块度下保持竞争力的同时提高了大图上的社区质量,并且扩展高效。消融实验证实,度校正和距离校准都是必不可少的。
英文摘要
Modularity is a standard objective for community detection, comparing observed intra-community connectivity with its null-model expectation. Modern node embeddings capture structural and semantic proximity beyond raw topology, but replacing adjacency with representation-induced affinity while retaining the classical degree-based null model makes the observed term geometry-dependent and the expected term geometry-agnostic. This creates a null-model mismatch: the objective may reward affinity already expected from representation proximity. This paper introduces RAM, a representation-aware modularity objective with a distance-calibrated null model. RAM preserves the observed-minus-expected principle by calibrating expected affinity with both degree heterogeneity and representation distance. We analyse its relation to classical modularity and develop divisive and multi-level greedy optimisers. Experiments on seven real-world datasets and synthetic scalability tests show that RAM improves over topology-only, direct embedding-clustering, and deep learning-based baselines, remains competitive under classical modularity while improving community quality on large graphs, and scales efficiently. Ablations confirm that both degree correction and distance calibration are essential.