AI 中文总结
本文针对异构信息网络中稠密P部子图搜索问题,提出BoxDPpS方法,通过优化搜索与剪枝策略降低成本,在7个真实数据集上实现平均27.04倍加速且保持精确最优解。
AI 中文摘要
异构信息网络(HIN)对类型化实体和类型化关系进行建模,其中稠密的跨类型结构可揭示凝聚性语义模式,例如高产的作者-论文-会议组。给定一个查询元路径,稠密P部子图搜索(DPpS)问题会在每个类型化位置联合选择一个非空顶点集,并最大化由所选集大小的几何平均值归一化的诱导元路径实例数量。现有精确方法通过在iRM集上搜索并将每个固定M问题归约为最小割计算来求解DPpS,但其可扩展性受限于候选iRM集数量庞大以及重复求解大型辅助网络的高成本。本文提出BoxDPpS,一种高效的精确方法,可减少这两类成本:它执行带安全区域剪枝的盒级搜索,消除同一iRM集的冗余表示,通过有界预热改进早期剪枝,并压缩每个固定M辅助网络以进行精确参数伪流求解。在7个真实世界数据集上的实验表明,BoxDPpS在保持DPpS精确最优解的同时,较现有最优方法实现了平均27.04倍的加速。
英文摘要
Heterogeneous information networks (HINs) model typed entities and typed relations, where dense cross-type structures can reveal cohesive semantic patterns such as prolific author-paper-venue groups. Given a query meta-path, the densest P-partite subgraph search (DPpS) problem jointly selects a nonempty vertex set at each typed position and maximizes the number of induced meta-path instances normalized by the geometric mean of the selected set sizes. Existing exact methods solve DPpS by searching over iRM-sets and reducing each fixed-M problem to minimum-cut computations. However, their scalability is limited by the large number of candidate iRM-sets and the high cost of repeatedly solving large auxiliary networks. In this paper, we propose BoxDPpS, an efficient exact approach that reduces both sources of cost. It performs box-level search with safe region pruning, eliminates redundant representations of the same iRM-set, improves early pruning through bounded warm-up, and compresses each fixed-M auxiliary network for exact parametric pseudoflow solving. Experiments on seven real-world datasets show that BoxDPpS preserves the exact DPpS optimum while achieving an average speedup of 27.04x over the state-of-the-art method.