arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36648cs.CV

VLM4Cluster:视觉语言预训练时代的深度聚类基准

VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training

  • University College London(伦敦大学学院)
  • The University of Queensland(昆士兰大学)
  • Southeast University(东南大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Auckland University of Technology(奥克兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuanwei Hu, Bo Peng, Yuheng Jia, Xinting Hu, Yadan Luo, Wenjie Zhu

AI总结:

针对语言辅助图像聚类研究缺乏统一基准的问题,提出VLM4Cluster,涵盖17种方法、20个数据集,从有效性、鲁棒性、泛化与效率多维度评估,发现LaIC显著提升性能但存在局限。

AI中文摘要:

视觉语言预训练重塑了图像聚类领域,催生了语言辅助图像聚类(LaIC),该方法利用文本语义来补充视觉表示。尽管LaIC方法迅速涌现,但LaIC究竟在多大程度上推动了图像聚类的发展仍不明确,因为现有研究普遍存在重大局限性,包括实验设置不一致、数据集选择不充分以及评估维度有限。为弥补这一空白,我们提出了VLM4Cluster,一个面向预训练视觉语言模型(VLM)时代图像聚类的综合基准。VLM4Cluster实现了涵盖经典、深度和语言辅助图像聚类的17种代表性方法,并在涵盖经典、挑战性、细粒度、大规模和分布外设置的20个数据集上对其进行了评估。除有效性外,VLM4Cluster还沿三个互补维度系统性地研究了图像聚类:对对抗扰动的鲁棒性、分布偏移下的泛化能力以及计算效率。我们的研究表明,LaIC在许多语义要求较高的基准上显著推进了聚类性能的前沿,通常在分布偏移下表现出更强的泛化能力,并实现了更优的有效性-效率权衡。然而,其增益在大规模和细粒度数据集上变得不太一致,而语言辅助并未系统性地降低对对抗扰动的敏感性。VLM4Cluster已在此https URL发布。

英文摘要:

Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as existing studies generally suffer from major limitations, including inconsistent experimental settings, inadequate dataset selection, and limited evaluation dimensions. To address this gap, we introduce VLM4Cluster, a comprehensive benchmark for image clustering in the era of pre-trained vision-language models (VLMs). VLM4Cluster implements 17 representative methods spanning classical, deep, and language-assisted image clustering, and evaluates them on 20 datasets covering classical, challenging, fine-grained, large-scale, and out-of-distribution settings. Beyond effectiveness, VLM4Cluster systematically investigates image clustering along three complementary dimensions: robustness to adversarial perturbations, generalization under distribution shifts, and computational efficiency. Our study shows that LaIC substantially advances the clustering performance frontier on many semantically demanding benchmarks, generally exhibits stronger generalization under distribution shifts, and achieves a more favorable effectiveness-efficiency trade-off. However, its gains become less consistent on large-scale and fine-grained datasets, while language assistance does not systematically reduce sensitivity to adversarial perturbations. VLM4Cluster is released at https://github.com/YuanweiHuu/VLM4Cluster.

↑