面向图像分类的视觉-语言模型半监督适配
Semi-Supervised Adaptation of Vision-Language Models for Image Classification
浏览论文内容
中文总结 AI 辅助
本研究针对视觉-语言模型适配遥感卫星图像时标注样本匮乏的问题,提出SE-CLIP半监督框架,通过双阶段流程与类平衡选择策略提升性能,在UCM、NWPU基准上优于现有方法。
中文摘要 AI 辅助
CLIP等视觉-语言模型在处理自然图像方面展现出巨大潜力,但其性能常受卫星图像独特特性的限制。虽存在参数高效的适配技术,但其效果常因标注样本匮乏而受限。本研究提出Self-Evolutionary CLIP(SE-CLIP),一种专为场景分类设计的半监督框架,用于递归标签挖掘。该方法遵循双阶段流程:先基于少量标注种子进行初始预热,随后进入递归发现阶段,从无标注池中迭代识别高置信度样本。为维持演化支持集的完整性,采用类平衡选择策略,防止模型被易学习类别主导。在UCM和NWPU基准上的结果表明,SE-CLIP显著优于现有半监督方法,该框架为以最少人工干预将视觉-语言模型(VLMs)适配到遥感领域提供了可行解决方案。
英文摘要
Vision-language models like CLIP have shown sig- nificant potential in handling natural images, yet their perfor- mance is often limited by the distinct characteristics of satellite imagery. While parameter-efficient adaptation techniques exist, their efficacy is frequently limited by the scarcity of annotated samples. In this letter, we propose Self-Evolutionary CLIP (SE- CLIP), a semi-supervised framework designed for recursive label mining in scene classification. The approach follows a dual-phase pipeline, where an initial warm-up on a few annotated seeds is followed by a recursive discovery phase that iteratively identifies high-confidence samples from unlabeled pools. To maintain the integrity of the evolving support set, we employ a class-balanced selection strategy that prevents the model from being dominated by easily learned categories. Results on the UCM and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches. The framework provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
发表机构
- Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)
- King Saud University(沙特国王大学)
- Tarim University(塔里木大学)
- Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。