发表机构
ICTEAM, UCLouvain(天主教鲁汶大学 ICTEAM 研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对小样本遥感场景分类问题,提出LC-TIM方法,融合多源遥感基础模型亲和图,建立首个直推式小样本遥感场景分类基准,实验表明其性能优于现有方法。
AI 中文摘要
遥感场景分类越来越依赖于在大规模地球观测数据上预训练的基础模型。此外,直推式推理利用整个未标记查询集的集体统计结构,似乎天然适配遥感流程,其中大图像通常被分割成图像块并作为一批进行推理。本研究提出LC-TIM(局部一致直推式信息最大化),它扩展了最先进的小样本CLIP直推式信息最大化(TIM++)目标,引入局部一致性正则化器,该正则化器强制每个查询样本与其特征空间中的κ个最近邻之间的预测一致性。该正则化器以单一乘法因子的形式加入闭式q更新,仅添加可忽略的计算开销。我们进一步提出多源扩展,融合来自多个遥感基础模型的亲和图,进一步提升分类准确率。为评估这些方法,我们建立了首个全面的开源直推式小样本遥感场景分类基准,在10个不同数据集、2个遥感视觉-语言模型及各种小样本设置下,对LP++、TransCLIP、TIM++和LC-TIM进行评估。实验表明,直推式方法始终优于零样本基线,且LC-TIM达到最先进的准确率,在邻域线索最具信息量的少样本场景中获得最大增益。代码公开可用:https://github.com/...(注:原文未提供完整链接,保留格式)
英文摘要
Remote sensing scene classification is increasingly relying on foundation models pre-trained on large-scale Earth-observation data. Moreover, transductive inference, which exploits the collective statistical structure of the entire unlabeled query set, appears to naturally match remote sensing pipelines where large images are routinely split into patches and inferred as a batch. In this work, we introduce LC-TIM (Locally Consistent Transductive Information Maximization), which extends the state-of-the-art Transductive Information Maximization for Few-Shot CLIP (TIM++) objective with a local consistency regularizer that enforces prediction agreement between each query sample and its $κ$ nearest feature-space neighbors. The regularizer enters as a single multiplicative factor in the closed-form $q$-update, adding negligible computational overhead. We further propose a multi-source extension that fuses the affinity graph from multiple remote sensing foundation model, further boosting classification accuracy. To assess these methods, we establish the first comprehensive, open-source benchmark for transductive few-shot RS scene classification, evaluating LP++, TransCLIP, TIM++, and LC-TIM across ten diverse datasets, two remote sensing vision-language models, and across various few-shot settings. Our experiments show that transductive methods consistently outperform zero-shot baselines, and that LC-TIM achieves state-of-the-art accuracy, with the largest gains in the low-shot regime where neighborhood cues are most informative. Code is publicly available at: https://github.com/elkhouryk/LC-TIM
CommentsAccepted at ECCVW2026