基于LLM符号化结构化过程的可解释无监督社区检测
Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes
浏览论文内容
中文总结 AI 辅助
本研究提出LLM引导的无监督社区检测方法LUCID,通过四阶段流程结合LLM生成规则实现可解释性,在真实数据集上性能优于主流无监督与半监督基线。
中文摘要 AI 辅助
社区检测是图分析中的一项基础任务,旨在识别具有相似行为或兴趣的内聚实体组。经典的目标驱动方法难以应对复杂的图结构,而深度学习方法虽提升了性能,却牺牲了可解释性,且依赖标注数据与训练。具备强推理能力和世界知识的大语言模型(LLM),有望用于可解释、无标注的社区检测。为利用这些优势,我们提出LUCID——一种LLM引导、可解释、免训练的无监督社区检测方法。受自然系统相变动力学启发(复杂结构通过初始化、合并、优化与选择产生),LUCID被设计为四阶段流程:1. 局部视图社区初始化阶段,利用k-ego上下文与无监督节点角色编码局部图结构;2. 多因素社区合并阶段,采用LLM生成的规则迭代合并局部社区;3. 多粒度社区优化阶段,并行应用LLM生成的由粗到细规则以减少边界噪声;4. 全局视图社区选择阶段,基于拓扑紧凑性与边界清晰度识别高质量社区。在真实世界数据集上的大量实验表明,作为无监督方法的LUCID达到了最优性能,且始终优于领先的无监督与半监督基线方法。
英文摘要
Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests. Classic objective-driven methods struggle with complex graph structures, while deep-learning approaches improve performance at the expense of interpretability and rely on labeled data and training. Large language models (LLMs), with strong reasoning capabilities and world knowledge, are promising for interpretable, label-free community detection. To leverage these strengths, we propose LUCID, an LLM-guided, interpretable, training-free, and unsupervised community detection method. Inspired by phase-transition kinetics in natural systems, where complex structures emerge through initialization, merging, refinement, and selection, LUCID is designed as a four-stage pipeline. Within this pipeline, the LLM induces formal rules that translate implicit knowledge into explicit and interpretable logical structures. Specifically, (1) the Local-View Community Initialization stage encodes local graph structures using k-ego contexts and unsupervised node roles; (2) the Multi-factor Community Merge stage uses LLM-induced rules to iteratively merge local communities; (3) the Multi-grain Community Refinement stage applies LLM-induced coarse-to-fine rules in parallel to reduce boundary noise; and (4) the Global-view Community Selection stage identifies high-quality communities based on topological compactness and boundary clarity. Extensive experiments on real-world datasets demonstrate that LUCID, as an unsupervised approach, achieves state-of-the-art performance and consistently outperforms leading unsupervised and semi-supervised baselines.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。