AI 中文总结
本研究针对完全图相关聚类问题,提出簇线性规划的近似对偶分离方法,结合新取整方案得到1.3865+ε近似算法,改进此前最优结果,将簇线性规划整性间隙限定在[4/3,1.3865]区间。
AI 中文摘要
我们针对完全图中的相关聚类问题给出了(1.3865+ε)-近似算法,改进了Cao等人(STOC'24)之前的最优近似因子(1.485+ε)。我们有两个独立的关键贡献:一是针对簇线性规划对偶的高效近似分离 oracle,二是一种新的取整方案。对偶分离任务是CCMinRatio问题:对于带符号顶点权重q,在满足q(S)>0的集合S上最小化cost(S)/q(S);其中cost(S)衡量将S作为簇时,任何聚类中归因于S的相关聚类分歧。我们给出了一个随机(1+ε)-近似算法,运行时间为2^{poly(1/ε)}(n+⟨q⟩)^{O(1)}。通过椭球法,这可得到分数簇线性规划最优值的(1+ε)-近似,以及恰好可行的原始和对偶解,可证明其值并提供每个实例的证书。该算法直接在原始实例上运行,无需全局预聚类:一种新的定位技术将搜索限制在小范围内,同时保留违规情况,之后弱正则性处理由此产生的密集二次最小化。我们的取整方案保留了Cao等人基于簇的过程,但使用了连续条件枢轴规则,其分析依赖于带显式权重的单一方差不等式,以及对所得多项式不等式的精确计算机辅助验证。这也将簇线性规划的整性间隙置于狭窄区间[4/3,1.3865]内。
英文摘要
We give a deterministic $(1.3865+ε)$-approximation for correlation clustering on complete graphs, improving the previous best factor of $1.485+ε$ of Cao et al. (STOC'24). Our first main contribution is an efficient weak separation oracle for the cluster-LP dual. Given signed vertex weights $q$, it either finds a set $S$ with $q(S)>cost(S)$ or certifies that $q/(1+ε)$ is dual feasible, where $cost(S)$ measures the correlation clustering disagreements attributed to $S$ in any clustering in which $S$ is a cluster. Via the ellipsoid method, this yields $(1+ε)$-approximate primal and dual solutions for the fractional cluster LP in deterministic time $2^{poly(1/ε)}n^{O(1)}$ . The separator works directly on the original instance, without global preclustering: a localization argument restricts the search to a small universe, where weak regularity handles the resulting dense quadratic minimization. Complementing this result, we prove that exact dual separation and cluster-LP optimization are NP-hard, even for complete unweighted instances. Our second main contribution is a new rounding of the cluster LP. It combines cluster-based rounding with a continuous conditional pivot rule, and its analysis relies on single variance inequality with explicit weights. This bounds the cluster-LP integrality gap by $1.3865$, near its known lower bound of $4/3$, and gives a per-instance primal--dual certificate of approximation. Finally, we extend our separator to obtain a $(1.92+ε)$-approximation for seeded correlation clustering (where each cluster may contain at most one seed from a prescribed seed set), and to bounded-weight instances, yielding a deterministic polynomial-time approximate implementation of the cluster-insertion primitive used in combinatorial correlation clustering.
Commentsv2: we added a deterministic localization lemma, a hardness proof, and applications