具有同质性的增长超图
Growing Hypergraphs with Homophily
AI总结:
研究增长超图模型,放宽边独立性假设,边形成受先前边和节点标签影响,推导度数分布幂律,用最大似然技术构建算法,通过期望最大化估计参数,展示模拟退火社区检测方法,凸显纳入边和标签依赖性的好处。
AI中文摘要:
有许多现存的超图模型,其交互由节点间基于属性的同质性控制,但大多假设边在节点参数条件下独立。放宽此假设,我们研究一种增长超图的机制模型,其中边的形成受先前边和二元节点标签影响。边在此模型中作为先前边的噪声副本形成,节点从一条边到下一条边的传递取决于标签间的多种同质性机制。这些机制在超图中产生可调的分类结构。我们推导了该模型中度数分布的幂律,并描述了边中标签联合分布的长期动态。我们的模型定义了带标签超图的似然性,可使用标准最大似然技术构建算法。我们通过(随机)期望最大化在合成数据和真实数据上估计模型参数。这些估计给出了经验多adic系统中同质性操作的统计原则描述。我们还展示了一种通过模拟退火进行社区检测的方法,虽计算成本高,但在合成数据和某些对基于边独立性假设的社区检测技术有挑战的经验数据集上取得了有竞争力的结果。我们的发现突出了在高阶建模和数据分析中纳入边和标签依赖性的好处,并指出了未来工作的几个方向。
英文摘要:
There are many extant models of hypergraphs with interactions governed by attribute-based homophily between nodes, but most assume independence between edges conditional on node parameters. Relaxing this assumption, we study a mechanistic model of growing hypergraphs in which edge formation is influenced by both previous edges and binary node labels. Edges form in this model as noisy copies of previous edges, where the transmission of nodes from one edge to the next depends on multiple homophilic mechanisms between labels. These homophilic mechanisms give rise to tunable assortative structure in the hypergraph. We derive a power law for the degree distribution in this model and describe the long-term dynamics of the joint distribution of labels contained in edges. Our model defines a likelihood over a labeled hypergraph, allowing us to use standard maximum-likelihood techniques to structure algorithms. We estimate the model parameters on synthetic and real data via (stochastic) expectation maximization. These estimates give statistically-principled descriptions of the operation of homophily in empirical polyadic systems. We also demonstrate an approach to community detection via simulated annealing which, though computationally expensive, achieves competitive results on both synthetic data and certain empirical data sets known to be challenging to community detection techniques based on the edge-independence assumption. Our findings highlight the benefits of incorporating edge- and label-dependence in higher-order modeling and data analysis, and point to several directions for future work.