发表机构
Wuhan University; Zhongguancun Academy(武汉大学; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了遥感视频数据集RSVideo-10K与评估基准RSVideo-Bench,发现现有视觉-语言模型在遥感视频理解上存在不足,提出强化学习框架RSVideo提升小目标时空聚焦能力,取得显著性能提升。
AI 中文摘要
遥感视频可实现对目标属性变化、短期活动及场景演化的实时观测,记录了孤立图像无法捕捉的运动、动作、交互及场景变化。现有模型主要针对单张图像或长时间范围的离散时间观测,目前仍缺乏用于评估视觉-语言模型在连续遥感视频理解任务上的统一评估设置。我们推出RSVideo-10K,这是一个包含10773个实例、147万帧、17.02小时素材的遥感视频数据集,涵盖无人机和卫星平台;其固定评估基准RSVideo-Bench包含2731个测试实例,评估遥感视频理解的两个互补维度:L1感知与L2推理,覆盖7个能力组、17项任务。评估显示,当前视觉-语言模型仍难以恢复微小局部证据、跟踪短暂状态、利用受场景约束的空间关系。基于该分析,我们进一步提出RSVideo,这是一个用于小目标时空聚焦的强化学习框架,可跨帧选择与问题相关的区域并抑制冗余背景标记;在26个开源视觉-语言模型中,RSVideo搭配InternVL3.5-14B时取得9.01%的最大绝对提升,搭配Qwen3.6-27B时达到40.63%的最高准确率。代码将在指定网址公开。
英文摘要
Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanning a long time range. However, a unified evaluation setting for assessing vision-language models on continuous remote-sensing video understanding remains lacking. We introduce RSVideo-10K, a remote-sensing video dataset comprising 10,773 instances, 1.47 million frames, and 17.02 hours of footage, containing both unmanned aerial vehicles and satellite platforms. Its fixed evaluation benchmark, RSVideo-Bench, contains 2,731 test instances and evaluates two complementary aspects of remote-sensing video understanding: L1 Perception and L2 Reasoning, spanning seven capability groups and 17 tasks. Evaluations show that current vision-language models still struggle to recover small local evidence, track short-lived states, and use scene-constrained spatial relations. Based on this analysis, we further propose RSVideo, a reinforcement learning framework for small-target spatiotemporal focusing that selects question-relevant regions across frames and suppresses redundant background tokens. RSVideo achieves a maximum absolute improvement of 9.01% with InternVL3.5-14B and attains the highest accuracy of 40.63% with Qwen3.6-27B across 26 open-source vision-language backbones. Codes will be available at https://github.com/HongjieZhou0329/RSVideo.