arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无人机-开放词汇视频实例分割:无人机也需要开放词汇视频实例分割

UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation

Mingyu Dou, Shi Qiu, Ming Hu, Yifan Chen, Zhe Sun

arXiv 2607.08075首次发表:更新:

AI 中文总结

研究针对无人机视频感知在开放场景中支持灵活查询和细粒度实例级动态理解不足的问题,提出无需训练的AeroTrack统一框架,构建AeroVIS评估基准,其实例化变体在无人机场景表现优异,还发布框架和基准以支持后续研究。

AI 中文摘要

无人机视频广泛应用于交通监测、城市管理和应急救援等领域。然而,现有的无人机视频感知主要依赖于预定义类别的框级定位和轨迹关联,难以在开放场景中同时支持灵活查询和细粒度实例级动态理解。为此,我们引入了无人机开放词汇视频实例分割(UAV-OVVIS)任务,该任务可根据开放词汇查询在无人机视频中发现目标,并输出具有全局一致标识的实例级分割轨迹。考虑到无人机场景中实例级注释的稀缺性,我们提出了AeroTrack,这是一个无需训练的统一框架。AeroTrack以周期性开放词汇检测、短片段掩码传播和跨片段身份统一为核心,重用现有的视觉基础模型来实现UAV-OVVIS。基于此框架,我们实例化了五个AeroTrack变体,并构建了AeroVIS,这是一个用于UAV-OVVIS的评估基准,包含9个无人机目标类别和8279条轨迹。实验表明,AeroTrack在无人机场景中显著优于现有的通用视频实例分割方法,并表现出强大的开放词汇鲁棒性和泛化能力。为支持未来的研究,我们将AeroTrack和AeroVIS作为UAV-OVVIS的统一框架和基准发布。

英文摘要

Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video perception is largely limited to box-level detection and tracking over predefined categories, making it difficult to jointly support flexible queries and fine-grained instance-level understanding of temporal dynamics in open scenarios. To this end, we introduce a new task, UAV Open-Vocabulary Video Instance Segmentation (UAV-OVVIS), which aims to discover targets in UAV videos according to open-vocabulary queries and output instance segmentation trajectories with globally consistent identities. Considering the scarcity of instance-level annotations in UAV scenarios, we propose AeroTrack, a training-free framework that coordinates existing visual foundation models to realize UAV-OVVIS. AeroTrack performs target discovery and segmentation through periodic open-vocabulary detection and short-segment mask propagation, and introduces Lifecycle-aware ID Association (LIA) to recover global identities under segment-wise inference. Based on this framework, we instantiate five feasible variants and construct AeroVIS, a UAV-OVVIS evaluation benchmark containing 9 UAV object categories and 8,279 trajectories. Experiments show that AeroTrack achieves better overall performance than the evaluated OV-VIS methods transferred to AeroVIS, while demonstrating good open-vocabulary transferability and dense-target handling capability in long UAV videos. The AeroTrack framework and the AeroVIS dataset will be open-sourced upon acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑