arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32570cs.LG

CAESAR:基于自主嵌入空间聚合式重组的聚类方法

CAESAR: Clustering via Autonomous Embedding-Space Agglomerative Reorganization

Ilan Bacry, Rémi Devaux, Antoine Jardin

首次发表
浏览论文内容

中文总结 AI 辅助

提出CAESAR方法,通过训练重组网络优化预训练嵌入空间,无需预设聚类数K,在文本和图像数据集上优于现有深度聚类方法。

中文摘要 AI 辅助

基于最近邻图操作的聚类算法,如FINCH(第一整数邻居聚类层次),严重依赖于所给嵌入空间的质量。然而,预训练的视觉和语言模型嵌入并未针对此目的进行优化。我们提出CAESAR,一种重组预训练嵌入空间的方法:训练一个重组网络,将互最近邻拉近,将非邻居推开,从而产生一个更适合聚类的重组嵌入空间。CAESAR还提供第二个主要优势:它从不要求聚类数K。这一点很重要,因为在现实的无监督设置中,K通常是未知的,发现K往往是问题的一部分,而大多数强聚类方法却将其作为输入。因此,我们设计整个CAESAR流程来推断K,而不是假设K已知。实验上,重组嵌入在文本和图像数据集上始终优于原始空间的聚类效果。由于少数同时推断K的深度聚类方法未发布其代码,我们除了与在同一嵌入空间上推断K的方法进行受控比较外,还与给定真实K的强深度聚类方法进行比较,赋予后者显著的oracle优势。即便如此,CAESAR在文本上优于所有方法,在最具挑战性的图像基准上取得最佳结果,并在其他基准上保持竞争力。因此,重组预训练嵌入成为对现实数据进行聚类的简单而强大的途径,其中类别重叠且聚类数未知。

英文摘要

Clustering algorithms that operate on nearest-neighbor graphs, such as FINCH (First Integer Neighbor Clustering Hierarchy), depend heavily on the quality of the embedding space they are given. However, pretrained vision and language model embeddings are not optimized for this purpose. We propose CAESAR, a method that reorganizes a pretrained embedding space: a reorganization network is trained to pull mutual nearest neighbors together and push non-neighbors apart, yielding a reorganized embedding space substantially better suited to clustering. CAESAR offers a second major advantage: it never requires the number of clusters $K$. This matters because in realistic unsupervised settings, $K$ is typically unknown and discovering it is often part of the problem, yet most strong clustering methods take it as an input. We therefore design the entire CAESAR pipeline to infer $K$ rather than assume it is known. Empirically, reorganizing the embeddings consistently improves clustering over the raw space on both text and image datasets. Since the few deep clustering methods that also infer $K$ do not release their code, we complement controlled comparisons with methods that infer $K$ on the same embedding space by comparisons with strong deep clustering methods that are given the true $K$, giving them a substantial oracle advantage. Even so, CAESAR outperforms all of them on text, achieves the best results on the most challenging image benchmark and remains competitive on the others. Reorganizing pretrained embeddings thus emerges as a simple and powerful route to clustering realistic data, where classes overlap and the number of clusters is unknown.

补充信息

↑