arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解十亿像素尺度的动态场景:无人机广域时空感知

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan

arXiv 2609.18210首次发表:更新:

发表机构

Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无人机十亿像素动态场景理解,提出超高分辨率数据集HARD及延迟感知指标s-HOTA,系统评估检测、跟踪与VQA任务,揭示流程延迟对性能的关键影响。

AI 中文摘要

无人机载成像已从百万像素传感器发展到十亿像素传感器,使空中感知从识别单个目标转向理解整个动态场景。我们将这一需求定义为广域时空场景理解(WSTU),它要求同时具备广域覆盖、目标级分辨率和时间连续性,而现有数据集缺乏这一组合。为填补这一空白,我们引入了超高分辨率(12768x9564)机载遥感数据集(HARD),该数据集在三个层面进行标注,用于目标检测、多目标跟踪和场景级视觉问答。超高分辨率图像将每帧处理时间提升至数秒。在此尺度下,评估中不能再忽略延迟。因此,我们提出了一种用于多目标跟踪的延迟感知指标,称为流式HOTA(s-HOTA)。广泛的基线实验表明,超高分辨率处理如何重塑每项任务。对于检测,端到端流程对精度和速度的影响与检测器本身一样大。对于跟踪,高延迟对关联轴的影响远比对检测轴的影响更不均匀,而关联正是流程产生分歧之处。因此,在离线状态下表现最佳的流程在s-HOTA下可能失去领先地位。对于VQA,视觉语言模型在跨帧身份绑定方面仍然薄弱,无法将其单帧优势迁移到该任务。这些发现共同表明,我们评估的基线未能达到WSTU的要求。HARD提供了数据和系统化基线以推进WSTU。

英文摘要

UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand as Wide-area Spatio-temporal Scene Understanding (WSTU), which requires wide-area coverage, per-target resolution, and temporal continuity at once, a combination existing datasets lack. To fill this gap, we introduce an ultra-High-resolution (12768x9564) Airborne Remote-sensing Dataset (HARD) annotated at three levels for object detection, multi-object tracking, and scene-level visual question answering. Ultra-high-resolution imagery raises per-frame processing time to seconds. At that scale latency can no longer be ignored in evaluation. Thus, we propose a latency-aware metric for multi-object tracking called streaming-HOTA (s-HOTA). Extensive baseline experiments show how ultra-high-resolution processing reshapes each task. For detection, the end-to-end pipeline affects accuracy and speed as much as the detector itself does. For tracking, high latency charges the association axis far more unevenly than the detection axis, and association is where pipelines diverge. As a result, the pipeline that performs best offline can lose its lead under s-HOTA. For VQA, vision-language models remain weak at cross-frame identity binding and cannot transfer their single-frame gains to it. Together these findings show that the baselines we evaluate fall short of WSTU. HARD provides the data and the systematic baselines to advance it.

Comments9 pages, 5 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑