arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从枚举到覆盖:大型异构信息网络上的近最优P部稠密子图搜索

From Enumeration to Covering: Near-Optimal Densest P-Partite Subgraph Search over Large Heterogeneous Information Networks

Lu Chen, Chengfei Liu, Rui Zhou, Jiajie Xu, Jianxin Li

arXiv 2608.06906首次发表:更新:

AI 中文总结

本文针对异构信息网络的稠密P部子图问题,提出用覆盖替代枚举的自适应剥离算法,实现近最优近似比,在真实网络上较枚举方法大幅加速,可验证子图近最优性。

AI 中文摘要

给定异构信息网络(Heterogeneous Information Network, HIN)和长度为i的查询元路径P,稠密P部子图问题旨在找到跨越P的i个类型层的子图,以最大化无参数密度:元路径实例数量除以各层大小的几何均值。该问题在文献、电商和生物医学网络中均有应用。现有最优近似方法通过固定每层权重将几何均值目标线性化,但需为每个可行权重集求解一个子问题,此类权重集数量为O((n/i)^i),且每个子问题仅能达到1/i的近似比。本文证明既无需穷举枚举也无需宽松近似比:首先,用覆盖替代枚举——仅需多对数级数量的代表性权重集,且通过数据相关边界进一步局部化,即可覆盖所有可行权重集,同时密度仅损失可调因子1+η;其次,将每个固定权重子问题转化为加权超模稠密子图实例并近最优求解,将整体近似比提升至(1-δ)/(1+η)。据本文所知,这是首个超越二分图(i=2)情形的近最优密度近似方法,且对任意固定i均给出PTAS。算法层面,求解器采用自适应剥离方案,无需显式生成元路径实例(其数量可远超图规模数个数量级);基于当前最优解的归约进一步在求解子问题前丢弃部分代表性权重集。在五个真实HIN上的实验表明,本文算法较基于枚举的基线方法实现了显著加速,且可进一步验证返回子图的近最优性。

英文摘要

Given a heterogeneous information network (HIN) and a query meta-path P of length i, the densest P-partite subgraph problem finds the subgraph, spanning the i typed layers of P, that maximizes a parameter-free density: the number of meta-path instances over the geometric mean of the layer sizes. It has applications across bibliographic, e-commerce, and biomedical networks. The state-of-the-art approximation linearizes the geometric-mean objective by fixing per-layer weights, but solves one subproblem for every feasible weight set, of which there are $O((n/i)^i)$, and on each achieves only a $1/i$ approximation. We show that neither the exhaustive enumeration nor the loose guarantee is necessary. First, we replace enumeration by covering: polylogarithmically many representative weight sets, localized further by a data-dependent bound, cover all feasible ones while losing only a tunable factor $1+η$ in density. Second, we cast each fixed-weight subproblem as a weighted supermodular densest-subgraph instance and solve it near-optimally, lifting the overall guarantee to $(1-δ)/(1+η)$. To our knowledge, this is the first near-optimal density approximation beyond the bipartite ($i=2$) case, and it yields a PTAS for every fixed i. Algorithmically, our solver is an adaptive peeling scheme that never materializes the meta-path instances, whose number can exceed the graph size by orders of magnitude. An incumbent-driven reduction further discards representative weight sets before their subproblems are solved. Experiments on five real HINs show that our algorithms achieve substantial speedups over enumeration-based baselines and can further certify the near-optimality of the returned subgraph.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑