arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35507cs.CV

ReVA:超越重复的遥感视频问答场景中心数据集

ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering

发表机构理海大学 · 南加州大学 · 高通人工智能研究院
查看机构详情
  • Lehigh University(理海大学)
  • University of Southern California(南加州大学)
  • Qualcomm AI Research(高通人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Zhen Yao, Likai Wang, Yuming Yang, Zhihao Zheng, Bo Lang, Qiuyu Tang, Jialu Sheng, Jingqi Xu, Yuehai Yang, Jumal Barker, Xiaowen Ying, Mooi Choo Chuah

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有遥感基准的重复问题和静态图像局限,本文提出ReVA数据集,包含2,438个无人机视频和22K问答对,评估MLLMs的时空推理能力,并揭示当前模型的不足。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在遥感领域展现了显著的进步。然而,现有的遥感多模态推理基准存在两个关键局限:它们依赖于(i)模板驱动的问题,导致问题重复;(ii)静态图像,无法捕捉无人机/UAV视频固有的时间特性。这使得遥感视频推理的系统性评估在很大程度上未被探索。为弥补这一空白,我们引入了ReVA,一个用于遥感视频问答的新数据集,旨在评估MLLMs的时空、场景中心和推理导向能力。ReVA包含覆盖全球18个城市的2,438个无人机视频(580K帧)和22K个高质量问答对,涵盖11项具有挑战性的QA任务。我们开发了一个半自动标注流程,利用文本LLMs和MLLMs生成问答对,并进行人工验证。我们在ReVA上评估了23个专有和开源的视频LLMs,揭示了当前模型的根本性局限。这些发现使ReVA成为推动更好遥感视频理解和时间推理能力以实现实际部署的关键基准。我们的代码和数据集可在以下网址获取:this https URL

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA

↑