arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

T-STAR:卫星视频中时空全景场景图生成的大规模基准测试

T-STAR: A Large-Scale Benchmark for Spatio-Temporal Panoptic Scene Graph Generation in Satellite Video

Linlin Wang, Xue Yang, Zhihuang Zhou, Zhenyu Zhong, Ruiyuan Zhang, Yansheng Li

arXiv 2607.21228首次发表:更新:

发表机构

School of Remote Sensing and Information Engineering, Wuhan University; School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(武汉大学遥感信息工程学院; 上海交通大学自动化与智能感知学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对卫星视频结构化理解难题,提出时空全景场景图生成任务,构建T-STAR大规模基准数据集,含超百万实例掩码和时空三元组,还提出统一框架,经实验验证其对卫星视频理解研究的重要性和框架有效性。

AI 中文摘要

对卫星视频进行结构化理解对于推进动态地理空间场景分析从低级感知到高级认知至关重要。本文引入卫星视频中的时空全景场景图生成(TPSG)作为一项新的基准任务,旨在生成由一组具有明确时间跨度的三元组<主体,关系,对象>组成的结构图,通过联合建模身份一致的实例掩码和全景场景元素之间的时空关系来描述动态地理空间场景。然而,目前尚无用于卫星视频TPSG的专用数据集,且卫星视频中的TPSG具有挑战性,自然视频的TPSG模型不适用于卫星视频。本文提出T-STAR,一个用于卫星视频TPSG的大规模基准数据集,包含超过110万个实例掩码和超过380万个时空三元组。为实现卫星视频TPSG,还提出统一框架增强跨帧实例一致性和时空关系预测。大量实验证明了T-STAR的重要性和所提框架的有效性,为未来结构化卫星视频理解研究建立了强大基准。数据集和代码可通过链接获取。

英文摘要

Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from low-level perception to high-level cognition. To move beyond object-centric perception, this paper introduces spatio-temporal panoptic scene graph generation (TPSG) in satellite video as a new benchmark task. TPSG aims to generate a structured graph composed of a set of triplets <subject, relationship, object> with explicit temporal spans, thereby describing dynamic geospatial scenes by jointly modeling identity-consistent instance masks and spatio-temporal relationships among panoptic scene elements. However, there is still no dedicated dataset for TPSG in satellite video. Moreover, TPSG in satellite video is intrinsically challenging, as objects are often small and weakly textured, cross-frame association is easily disrupted by occlusion and background clutter, and relationship semantics are highly coupled with spatial structure and temporal evolution. Consequently, TPSG models developed for natural videos are not directly applicable to satellite video. This paper presents T-STAR, a large-scale benchmark dataset for TPSG in satellite video, comprising over 1.1 million instance masks and over 3.8 million spatio-temporal triplets across 39 fine-grained object categories and 70 fine-grained relationship categories. To enable TPSG in satellite video, we propose a unified framework to enhance cross-frame instance consistency and spatio-temporal relationship prediction. Extensive experiments demonstrate the significance of T-STAR and the effectiveness of the proposed framework, establishing a strong benchmark for future research on structured satellite video understanding. The dataset and code are available at https://github.com/linlin-dev/T-STAR.

Comments17 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑