带约束的相关聚类:迈向最优性
Constrained Correlation Clustering: Towards Optimality
浏览论文内容
中文总结 AI 辅助
本研究针对带约束的相关聚类,将其最优近似因子区间缩小至[2,16/7-γ),并证明聚类删除的2-近似算法是最优的,明确了其与普通相关聚类的近似难度差距。
中文摘要 AI 辅助
在相关聚类问题中,给定一个无向图,任务是计算一个聚类(节点的划分),使违反对的数量最少(即不同聚类间的边数加上同一聚类内非边的数量)。该问题的带约束版本中,目标是计算一个满足额外硬约束的聚类,这些约束要求某些对必须在同一聚类,某些对必须在不同聚类。本研究将带约束的相关聚类(Constrained Correlation Clustering)确定为相关聚类的一个变体,其最优近似解有望实现,并在该方向取得进展。带约束的相关聚类是APX难问题,已知其最优近似因子位于区间(24/23, 3]内。我们大幅缩小了该区间,证明最优近似因子位于[2, 16/7 - γ),其中γ为一个小的正常数。我们的2的下界表明,相关聚类(其近似比可达1.485+ε)与带约束的相关聚类之间存在差距。我们的16/7 - γ的上界使用了Sherali-Adams松弛,且超出了基于三角形的直接分析;更准确地说,我们的算法属于一类自然的枢轴算法,我们证明这类算法无法通过基于三角形的直接分析得到优于16/7的近似比。最后,作为我们技术的副产品,我们完全解决了聚类删除(Cluster Deletion)的近似性问题。聚类删除是带约束的相关聚类中一个被广泛研究的特例,已知存在一个2-近似算法。我们证明该算法是最优的,因为我们的下界甚至对该特例也成立。
英文摘要
In the Correlation Clustering problem, we are given an undirected graph and are tasked with computing a clustering (partition of the nodes) that minimizes the number of violated pairs (edges across different clusters plus non-edges within clusters). In the constrained version of this problem, the goal is to compute a clustering that satisfies additional hard constraints mandating certain pairs to be in the same cluster and certain pairs to be in different clusters. In this work, we identify Constrained Correlation Clustering as a variant of Correlation Clustering for which optimal approximations might be within reach, and make progress towards this front. Constrained Correlation Clustering is APX-Hard, and the optimal approximation factor is known to lie in $(\frac{24}{23},3]$. We significantly tighten this gap, by showing that the optimal approximation factor lies in $[2,\frac{16}{7}-γ)$ for a small constant $γ>0$. Our lower bound of $2$ shows a separation between Correlation Clustering (which admits an $1.485+ε$ approximation) and Constrained Correlation Clustering\footnote{The same hardness result was obtained independently by Cao and Xu~\cite{cao2026clusterdeletionhardapproximate}.}. Our upper bound of $\frac{16}{7}-γ$ uses the Sherali-Adams relaxation and goes beyond straightforward Triangle-Based analysis; more precisely, our algorithm belongs to a natural class of pivoting algorithms for which we prove that a straightforward Triangle-Based analysis cannot prove a better-than-$\frac{16}{7}$ approximation. Finally, as a byproduct of our techniques, we completely resolve the approximability of Cluster Deletion. Cluster Deletion is a well-studied special case of Constrained Correlation Clustering for which a $2$-approximation algorithm is known. We show that this is optimal, as our lower bound holds even for this special case.