arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于软社区结构的可扩展分层图生成

Scalable Hierarchical Graph Generation via Soft Community Structure

Ahmet Tüzen, Helge Langseth, Kjetil Nørvåg

arXiv 2610.12163首次发表:更新:

发表机构

Norwegian University of Science and Technology(挪威科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Schema模型,通过递归分解参考图为软社区层次结构,分三阶段独立训练生成大型属性图,在四个真实图上表现更优,且在千万级节点图上具备可扩展性。

AI 中文摘要

生成大型属性图需要复现其拓扑结构、联合生成属性与结构,同时保持可扩展性。许多真实世界的图表现为单个大图,因此生成模型必须从其拟合的单个图中泛化,而无需独立样本。我们提出Schema,它将参考图递归分解为软社区的层次结构,为每个节点分配成员分布。生成过程分为三个独立训练的阶段:(1)基于软成员关系合成节点属性;(2)从局部结构上下文生成社区内边;(3)对桥节点上的社区间连接建模,桥节点的成员质量分布在多个社区中。各阶段均不形成完整邻接矩阵,且每个阶段仅在社区大小限定的子图上运行。我们还引入了涵盖结构保真度、记忆性、下游效用和可扩展性的评估协议。在四个真实世界属性图上,Schema在复现参考图边的一小部分的同时,比任何其他生成属性的模型更紧密地恢复了局部与长程结构之间的平衡。它保留了参考图的下游准确性,未将其人为提高到该水平之上。与其结构保真度匹配的基线会记忆参考图,而具有更高下游准确性的基线要么超过参考准确性,要么无法在更大的图上完成。我们在另外六个节点数达1000万的图上测量了可扩展性。

英文摘要

Generating large attributed graphs requires reproducing the topology, generating attributes jointly with the structure, and remaining scalable. Many real-world graphs exist as a single large graph, so a generative model has to generalize from the one graph it is fit on, without independent samples. We present Schema, which recursively decomposes a reference graph into a hierarchy of soft communities, assigning each node a membership distribution. Generation is then split into three stages, each trained independently: (1) synthesizing node attributes conditioned on soft memberships, (2) generating intra-community edges from local structural context, and (3) modeling inter-community connections over bridge nodes whose membership mass is distributed across several communities. No stage forms the full adjacency matrix, and each stage operates on a subgraph bounded by the community size. We also introduce an evaluation protocol that covers structural fidelity, memorization, downstream utility, and scalability. On four real-world attributed graphs, Schema recovers the balance between local and long-range structure more closely than any other model that generates attributes, while reproducing only a small fraction of the reference edges. It retains the downstream accuracy of the reference graph without raising it artificially above that level. Baselines that match its structural fidelity memorize the reference, while those with higher downstream accuracy either exceed the reference accuracy or fail to complete on the larger graphs. We measure scalability on six additional graphs with up to 10 million nodes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑