发表机构
School of Computer Science and Engineering, Tianjin University of Technology; Engineering Research Center of Learning-Based Intelligent System, Ministry of Education of the People’s Republic of China, Tianjin University of Technology; School of Artificial Intelligence, Tianjin University; Engineering Research Center of City Intelligence and Digital Governance, Ministry of Education of the People’s Republic of China, Tianjin University(天津理工大学计算机科学与工程学院; 中华人民共和国教育部基于学习的智能系统工程研究中心,天津理工大学; 天津大学人工智能学院; 中华人民共和国教育部城市智能与数字治理工程研究中心,天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在无源数据时将预训练源模型适应到目标域的问题,提出LFM框架,利用视觉语言模型计算相似度确定标签转移类型、识别未知样本,通过共识策略精炼伪标签训练目标模型,实验验证了该框架的有效性和优越性。
AI 中文摘要
无源通用域适应(SF-UniDA)在协变量和标签转移情况下,在无法访问源数据时将预训练源模型适应到未标记目标域。现有方法依赖低效技术。基础模型在SF-UniDA中未充分探索。本文提出LFM框架,用视觉语言模型计算目标样本与文本标签相似度,确定标签转移类型,识别未知样本。通过共识策略,用预训练源模型初始化目标模型来精炼伪标签,并用其训练目标模型。实验证明了LFM框架的有效性和优越性。
英文摘要
Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a framework that leverages foundation models (LFM) for SF-UniDA. We use a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model. The label shift type is determined by analyzing the coefficient of variation of a similarity-based sample-level score. Unknown samples are identified using a binary Gaussian mixture model fitted to another similarity-based metric. Under a consensus strategy, the pseudo-labels generated by the VLM are refined by the target model initialized with the pre-trained source model, integrating knowledge from both the source domain and foundation models. Finally, these refined pseudo-labels are used to train the target model. Extensive experiments across all possible label shifts and multiple benchmarks demonstrate the effectiveness and superiority of our proposed LFM framework. Our code is available at https://github.com/iamjingli/LFM.
CommentsAccepted by IEEE Transactions on Multimedia (2026)