发表机构
MBZUAI Institute of Foundation Models; University of California, Los Angeles(穆罕默德·本·扎耶德人工智能大学基础模型研究所; 加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究有限超图中的最大强独立集问题,提出关联结构工具包,包括精确约简、上界、证书及分层贪心算法,并给出正确性与最优性保证,应用于多频带LSH-MinHash去重。
AI 中文摘要
我们研究有限超图中的最大强独立集问题:寻找最大的顶点集,使得该集合与每条超边的交集至多包含一个顶点。当每个观测到的块是局部不相容约束,但跨重叠块的传递闭包不合理时,就会出现这一目标。一个激励性示例是多频带LSH-MinHash去重,其中每个碰撞桶提供局部证据,而连通分量收缩可能强加虚假的全局等价关系。本文针对该问题开发了一套关联结构工具包。我们证明了支配、关联孪生和权重为1的块的精确约简;推导了闭式上界和低权重上界;引入了穿刺和覆盖证书以锐化这些界;并分析了一种由块权重和残差关联驱动的分层贪心聚类算法。算法分析包括可行性、极大性、条件最优性、分层见证匹配上界以及关联局部复杂度界。结果为广泛的关联族提供了正确性、终止性、不动点和最优性证书,并附有示例展示不同证书何时分离或重合。
英文摘要
We study the maximum strong independent set problem in a finite hypergraph: find the largest vertex set that intersects every hyperedge in at most one vertex. This objective arises whenever each observed block is a local incompatibility constraint but transitive closure across overlapping blocks is not justified. A motivating example is multi-band LSH-MinHash deduplication, where each collision bucket gives local evidence, while connected-component contraction can impose spurious global equivalences. The paper develops an incidence-structural toolkit for this problem. We prove exact reductions for dominance, incidence twins, and weight-1 blocks; derive closed-form and low-weight upper bounds; introduce puncturing and covering certificates that sharpen those bounds; and analyze a layered greedy clustering algorithm driven by block weights and residual incidence. The algorithmic analysis includes feasibility, maximality, conditional optimality, a layered witness-matching upper bound, and incidence-local complexity bounds. The results give correctness, termination, fixed-point, and optimality certificates for broad incidence families, together with examples showing when different certificates separate or coincide.