arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于结合文本信息的图像聚类的深度模态共享自表达学习

Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information

Xianghan Meng, Wei He, Zhiyuan Huang, Chun-Guang Li

arXiv 2608.08418首次发表:更新:

AI 中文总结

提出DeepMORSE方法,通过模态共享自表达模型学习结合文本信息的图像聚类,在6个基准数据集上聚类性能提升超3%,且所学表示可迁移至下游任务。

AI 中文摘要

利用文本信息进行图像聚类已成为一个有前景的研究方向,这在很大程度上得益于视觉-语言模型(Vision-Language Models, VLMs)所学习到的强大表示。现有方法通常为每张图像检索对应的文本配对,然后通过直接强化跨模态一致性来优化多模态表示,例如最大化预训练VLMs所继承的图像-文本相似度。然而,这种策略在对齐跨模态的异构表示时,没有显式建模每个模态内部的固有结构,因此可能产生不可靠的对齐结果,或破坏对聚类至关重要的模态特定结构。在本文中,我们提出了一种简单但有原则的方法,称为深度模态共享自表达模型(deep modality-shared self-expressive model, DeepMORSE),该方法通过模态共享自表达模型发现跨模态结构,同时学习符合模态特定子空间并集的结构化表示。此外,我们从理论上证明,模态共享自表达系数会抑制类间噪声,从而得到保留子空间的解决方案,并且还表明小批量优化过程会为自表达模型引入隐式正则化。我们在六个广泛使用的图像聚类基准上评估了DeepMORSE,观察到在UCF-101、DTD-47和ImageNet-Dogs数据集上的性能提升超过3%。此外,我们通过在图像检索、零样本分类等下游任务上实现了最先进的性能,证明了所学表示的强可迁移性——无需任何特定任务的损失或后处理。代码可在以下网址获取:this https URL

英文摘要

Leveraging textual information for image clustering has emerged as a promising direction, largely owing to the powerful representations learned by Vision-Language Models (VLMs). Existing approaches typically retrieve a textual counterpart for each image and then refine multimodal representations by directly enforcing cross-modal agreement, e.g., maximizing image-text similarity inherited from pretrained VLMs. However, such a strategy aligns heterogeneous representations across modalities without explicitly modeling the intrinsic structure within each modality and thus might yield unreliable alignment or distort modality-specific structures that are crucial for clustering. In this paper, we propose a simple but principled approach, termed deep modality-shared self-expressive model (DeepMORSE), which discovers cross-modal structures via a modality-shared self-expressive model and simultaneously learns structured representations that conform to a union of modality-specific subspaces. Moreover, we theoretically justify that the modality-shared self-expressive coefficients suppress inter-class noise towards a subspace-preserving solution, and show that mini-batch optimization procedure introduces an implicit regularization onto the self-expressive model. We evaluate our DeepMORSE on six widely used image clustering benchmarks and observe performance improvements exceeding 3% on the UCF-101, DTD-47, and ImageNet-Dogs datasets. In addition, we demonstrate the strong transferability of the learned representations by achieving state-of-the-art performance on downstream tasks such as image retrieval and zero-shot classification---without requiring any task-specific losses or post-processing. The code is available at: https://github.com/mengxianghan123/DeepMORSE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑