arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25485cs.CV

面向图像分类的视觉-语言模型半监督适配

Semi-Supervised Adaptation of Vision-Language Models for Image Classification

Mohamed L. Mekhalfi, Mohamad M. Al Rahhal, Yakoub Bazi, Salah E. Khenfer, Mingdeng Shi, Hua Zou, Mansour Zuair

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对视觉-语言模型适配遥感卫星图像时标注样本匮乏的问题,提出SE-CLIP半监督框架,通过双阶段流程与类平衡选择策略提升性能,在UCM、NWPU基准上优于现有方法。

中文摘要 AI 辅助

CLIP等视觉-语言模型在处理自然图像方面展现出巨大潜力,但其性能常受卫星图像独特特性的限制。虽存在参数高效的适配技术,但其效果常因标注样本匮乏而受限。本研究提出Self-Evolutionary CLIP(SE-CLIP),一种专为场景分类设计的半监督框架,用于递归标签挖掘。该方法遵循双阶段流程:先基于少量标注种子进行初始预热,随后进入递归发现阶段,从无标注池中迭代识别高置信度样本。为维持演化支持集的完整性,采用类平衡选择策略,防止模型被易学习类别主导。在UCM和NWPU基准上的结果表明,SE-CLIP显著优于现有半监督方法,该框架为以最少人工干预将视觉-语言模型(VLMs)适配到遥感领域提供了可行解决方案。

英文摘要

Vision-language models like CLIP have shown sig- nificant potential in handling natural images, yet their perfor- mance is often limited by the distinct characteristics of satellite imagery. While parameter-efficient adaptation techniques exist, their efficacy is frequently limited by the scarcity of annotated samples. In this letter, we propose Self-Evolutionary CLIP (SE- CLIP), a semi-supervised framework designed for recursive label mining in scene classification. The approach follows a dual-phase pipeline, where an initial warm-up on a few annotated seeds is followed by a recursive discovery phase that iteratively identifies high-confidence samples from unlabeled pools. To maintain the integrity of the evolving support set, we employ a class-balanced selection strategy that prevents the model from being dominated by easily learned categories. Results on the UCM and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches. The framework provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.

发表机构

  • Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)
  • King Saud University(沙特国王大学)
  • Tarim University(塔里木大学)
  • Wuhan University(武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

↑