arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38603cs.CV

Aperture:面向遥感的无训练多尺度概念瓶颈

Aperture: Training-Free Multiscale Concept Bottlenecks for Remote Sensing

Rishabh Mondal, Nipun Batra, Utkarsh Mall

首次发表
浏览论文内容

中文总结 AI 辅助

针对遥感模型缺乏可解释性的问题,提出无训练多尺度概念瓶颈方法APERTURE,结合贪心四叉树路由与预训练MLLM获取概念分数,在细粒度数据集SiFC上宏F1分数超越最佳无训练基线超10个百分点,并优于有监督概念瓶颈模型。

中文摘要 AI 辅助

尽管地球观测模型已取得显著进展,但它们仍缺乏可解释性。概念瓶颈模型虽然提供了可解释性和专家交互能力,但在遥感领域要么训练成本过高,要么在无标注情况下性能不佳。我们认为,在遥感等专家领域中,此类无训练模型需要在图像空间和概念空间中都具备精细细节。在图像空间,我们提出了一种使用贪心四叉树路由的多尺度概念瓶颈,以定位小尺度概念。在概念空间,我们用预训练的MLLM替代对比视觉语言模型,并提出了一种从这些模型中获取可靠概念分数的方法。我们引入了APERTURE,它在全局图像和原生概念尺度层面融合概念分数,以实现最先进的无训练模型性能。为测试这些模型,我们引入了SiFC,这是一个覆盖三个国家的细粒度概念中心数据集,具有人工审核的类别级概念图。在SiFC上,APERTURE在宏F1分数上比最佳无训练基线高出超过10个百分点,并且值得注意的是,它也优于有监督的概念瓶颈模型。针对性的组件移除测试检验了概念分数是否响应视觉证据的变化,而时间实验表明,描述符更新无需重新训练即可改善对技术变化的识别。

英文摘要

While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability and expert interaction, they are either too expensive to train for the remote sensing domain or perform poorly without annotation. We posit that in expert domains like remote sensing, such training-free models require both fine details in both image and concept space. In image space, we propose a multiscale concept bottleneck using greedy quadtree routing to locate small concepts. In concept space, we replace contrastive vision language models with pre-trained MLLMs and present a way to get reliable concept scores from them. We introduce APERTURE that blends concept scores at the global image and native concept-scale level to give state-ofthe-art training-free model performance. To test these models, introduce SiFC, a fine-grained concept-centric dataset across three countries, with human-reviewed class-level concept maps. On SiFC, APERTURE outperforms the best training-free baselines by more than 10 percentage points in macro F1-score, and notably also outperforms supervised concept bottleneck models. Targeted component-removal tests examine whether concept scores respond to changes in visual evidence, while temporal experiments show that descriptor updates improve recognition of technological changes without retraining.

发表机构

  • Indian Institute of Technology Gandhinagar(印度理工学院甘地讷格尔分校)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

↑