发表机构
Vrije Universiteit Amsterdam; CWI Amsterdam; UC Berkeley(阿姆斯特丹自由大学; 阿姆斯特丹数学与计算机科学研究中心; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对关联聚类中完全图与一般非完全图的近似保证差距,研究随机删边的不完全图,证明其近似保证优于一般图,算法近似比接近完全图。
AI 中文摘要
关联聚类是一类基础的无监督学习问题。在完全图上,最小分歧和最小最大目标均存在常数因子近似算法;但在一般(非完全)图上,最优保证分别膨胀至$O(\text{log }n)$和$O(\text{sqrt }n)$。两种场景的差距催生了如下问题:是否存在一类不完整图,可规避一般图的下界,且能达到接近完全图的近似保证?我们研究通过随机子采样完全带符号图$G$得到的自然图类,其中每条边以概率$q$独立删除。针对这类图实例,对于最小最大和最小分歧目标,我们证明了依赖于$q$的近似保证,其显著优于一般图可达到的界。我们将理论结果与实验结合,实验也表明我们算法的近似比接近完全图的近似比,且优于一般(非完全)图的最坏情况界。
英文摘要
Correlation clustering is a fundamental unsupervised learning problem. On complete graphs, both the min-disagreement and min-max objectives admit constant-factor approximations, yet on general (non-complete) graphs, the best guarantees blow up to $O(\log n)$ and $O(\sqrt{n})$. This gap between the two regimes motivates the following question: are there classes of incomplete graphs that circumvent the lower bounds on general graphs and admit approximation guarantees approaching those attainable on complete graphs? We study a natural class of graphs obtained by randomly subsampling a complete signed graph $G$, where each edge is independently deleted with probability $q$. For such graph instances both for the min-max and the min-disagreement objectives, we prove approximation guarantees (depending on $q$) that are substantially better than the bounds achievable for general graphs. We supplement our theoretical results with experiments that also suggest that the approximation ratios of our algorithm are close to those of the complete graph and better than the worst-case bounds for general (non-complete) graphs.
Comments17 pages, 5 figures, 6 tables