发表机构
Federal University of São Carlos (UFSCar)(圣卡洛斯联邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对主动提示学习中的冷启动问题,提出无监督迁移聚类与选择性查询(UTC+SQ)框架,利用无监督迁移模型生成伪标签进行语义聚类,提升查询代表性,在8个数据集中6个获得准确率提升。
AI 中文摘要
视觉-语言模型(VLMs)通过对齐视觉和文本表示能够实现令人印象深刻的零样本分类性能,但每个新任务仍然需要手工设计的提示。主动提示学习(APL)将主动学习(AL)和提示学习(PL)结合到一个框架中,利用VLM的先验知识迭代查询最具信息量的图像进行标注。然而,冷启动问题对APL方法仍然存在,即初始查询的性能可能比随机采样更差。虽然最近的最先进的APL方法通过平衡采样和多模态特征来缓解这一问题,但它们依赖基于距离的刚性聚类来对这些特征进行分组。这种简单的方法可能难以捕捉VLM固有的复杂高维语义分布,导致查询代表性欠佳。无监督迁移可以作为一种可能的替代方案应用于这些特征,因为它能够在没有任何监督的情况下推断任务的潜在人工标注。这样,样本可以被分组到语义连贯的簇中。本文提出了无监督迁移聚类与选择性查询(UTC+SQ)框架,该框架通过利用最先进的无监督迁移模型增强了最近的APL方法。这些模型生成高保真伪标签,建立语义有意义的簇,从而能够选择更相关的样本。实验评估表明,从基于距离的聚类转向基于投影的聚类提高了查询子集的代表性,在测试的8个数据集中有6个实现了准确率提升。
英文摘要
Vision-Language Models (VLMs) are able to achieve impressive zero-shot classification performance by aligning visual and textual representations, but each new task still demands handcrafted prompts. Active Prompt Learning (APL) combines Active Learning (AL) and Prompt Learning (PL) into a single framework, allowing for the usage of the VLM prior knowledge for iteratively querying the most informative images to be labeled. However, the cold-start problem is still relevant for APL methods, where the performance of the initial query can be worse than random sampling. While recent state-of-the-art APL methods mitigate with balanced sampling and multimodal features, they rely on rigid, distance-based clustering to group these features. This simplistic approach can struggle to capture the complex, high-dimensional semantic distributions inherent to VLMs, leading to suboptimal query representativeness. Unsupervised transfer can be applied to these features as a possible alternative, since it is capable of inferring the underlying human labeling of a task without any form of supervision. This way, samples can be grouped in semantically coherent clusters. This paper proposes Unsupervised Transfer Clustering with Selective Querying (UTC+SQ), a framework that enhances a recent APL approach by leveraging state-of-the-art unsupervised transfer model. These models generate high-fidelity pseudo-labels that establish semantically meaningful clusters, allowing for the selection of more relevant samples. Experimental evaluations demonstrate that shifting from distance-based to projection-based clustering improves the representativeness of the queried subset, achieving accuracy gains in 6 of the 8 datasets tested.
Journal ref39th SIBGRAPI - Conference on Graphics, Patterns, and Images (SIBGRAPI'26), 2026, pp. 1-6