arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08539cs.CV

RSJEV:基于多模态大语言模型的判别式遥感场景分类

RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models

Dongchen Si, Di Wang, Mingzhen Xu, Jing Zhang, Bo Du, Liangpei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

RSJEV提出一种单次通过的多模态判别决策框架,将遥感场景分类重构为候选条件化决策,消除自回归解码,在多个基准上以0.8B参数实现更优精度与效率。

中文摘要 AI 辅助

遥感场景分类是地球观测和地理空间分析中的一项基础任务。现有方法主要遵循三种范式:特定任务的视觉分类、视觉-语言相似性匹配和自回归多模态生成。然而,视觉分类器依赖于预定义的标签空间,基于CLIP的方法通过静态图像-文本对齐进行识别,而多模态大语言模型(MLLMs)对于具有明确候选类别的分类任务引入了不必要的令牌级生成。为了解决这些局限性,我们提出了RSJEV,一种用于遥感场景分类的单次通过多模态决策框架。与将分类表述为自回归文本生成的传统MLLMs不同,RSJEV将场景分类重新表述为候选条件化的多模态判别决策过程,其中视觉表示、任务指令和候选类别语义被联合建模。具体来说,我们引入了一个OnePass Decider,它提取多模态决策状态并直接在候选类别空间内估计类别概率,消除了自回归解码,同时保留了视觉-语言交互。在三个广泛使用的遥感场景分类基准(包括UC Merced、AID和NWPU-RESISC45)上进行的大量实验表明,与代表性的基于CNN、Transformer、Mamba、CLIP和MLLM的方法相比,RSJEV实现了优越的分类性能。此外,RSJEV显著降低了推理成本,并且仅用紧凑的0.8B参数模型就实现了更好的精度-效率权衡。这些结果证明了状态条件化的多模态决策对于高效遥感图像理解的有效性。代码将在https URL上提供。

英文摘要

Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradigms: task-specific visual classification, vision-language similarity matching, and autoregressive multimodal generation. However, visual classifiers rely on predefined label spaces, CLIP-based methods perform recognition through static image-text alignment, and multimodal large language models (MLLMs) introduce unnecessary token-level generation for classification tasks with explicit candidate categories. To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification. Unlike conventional MLLMs that formulate classification as autoregressive text generation, RSJEV reformulates scene classification as a candidate-conditioned multimodal discriminative decision process, where visual representations, task instructions, and candidate category semantics are jointly modeled. Specifically, we introduce a OnePass Decider that extracts multimodal decision states and directly estimates category probabilities within the candidate category space, eliminating autoregressive decoding while preserving vision-language interactions. Extensive experiments on three widely used remote sensing scene classification benchmarks, including UC Merced, AID, and NWPU-RESISC45, demonstrate that RSJEV achieves superior classification performance compared with representative CNN-, Transformer-, Mamba-, CLIP-, and MLLM-based methods. Moreover, RSJEV significantly reduces inference costs and achieves a better accuracy-efficiency trade-off with only a compact 0.8B-parameter model. These results demonstrate the effectiveness of state-conditioned multimodal decision making for efficient remote sensing image understanding. The code will be available at https://github.com/Dongtcs/RSJEV.

发表机构

  • Wuhan University(武汉大学)
  • State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(武汉大学测绘遥感信息工程国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑