arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越外观:移动空地平台上行人关联与跟踪的多线索框架及大规模基准测试

Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms

Ruiqi Wu, Bingliang Jiao, Ruize Han, Hangzheng Yu, Xunkai Jiang, Shining Wang, Yuanqi Hu, Wenxuan Wang, Peng Wang

arXiv 2607.23803首次发表:更新:

发表机构

School of Computer Science, Northwestern Polytechnical University; Ningbo Institute, Northwestern Polytechnical University; National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology; Shenzhen University of Advanced Technology(西北工业大学计算机科学学院; 西北工业大学宁波研究院; 国家空天地海一体化大数据应用技术工程实验室; 深圳先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对移动空地平台上行人关联与跟踪面临的视点变化问题,提出FUSION框架,含MAC和OMFS模块,还引入大规模基准测试RealMvMoAT,实验表明FUSION在多个测试中达领先性能,为相关研究提供了资源。

AI 中文摘要

多视图多目标关联与跟踪(MvMoAT)可跨摄像机视图关联对象并随时间跟踪,支持多平台协作感知中的身份持久性和法医轨迹重建。与传统多目标跟踪不同,MvMoAT面临频繁的视点变化,会扭曲外观并破坏跨视图关联和时间跟踪。我们提出了FUSION,一种用于多视图关联和识别的视点鲁棒特征统一框架。其多线索自适应组合(MAC)模块将视点不变线索与外观特征自适应集成,以改善跨视图关联,而在线多视图特征同步(OMFS)跨历史和跨视图帧聚合行人特征以进行时间一致的跟踪。我们还引入了RealMvMoAT,一个具有大量摄像机间和摄像内视点变化的大规模基准测试。它包含来自10个场景中7台摄像机(5台无人机和2台地面视图)的504.9K帧,有超过730万个带身份标签的边界框。所有摄像机都有随机且大幅度的运动。据我们所知,RealMvMoAT是迄今为止最大的MvMoAT数据集。其实验表明FUSION在RealMvMoAT和六个公共基准测试中取得了领先性能。

英文摘要

Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory reconstruction in multi-platform cooperative perception. Unlike conventional multiple object tracking, MvMoAT faces frequent viewpoint shifts that distort appearance and undermine cross-view association and temporal tracking. We propose FUSION, a viewpoint-robust Feature Unification framework for multi-view aSsociation and IdentificatiON. Its Multi-cue Adaptive Combination (MAC) module adaptively integrates viewpoint-invariant cues with appearance features to improve cross-view association, while Online Multi-view Feature Synchronization (OMFS) aggregates pedestrian features across historical and cross-view frames for temporally consistent tracking. We also introduce RealMvMoAT, a large-scale benchmark featuring substantial inter- and intra-camera viewpoint variation. It contains 504.9K frames from 7 cameras (5 UAV and 2 ground views) across 10 scenes, with over 7.3M identity-labeled bounding boxes. All cameras exhibit random and substantial motion. To the best of our knowledge, RealMvMoAT is the largest MvMoAT dataset to date. Its scale, viewpoint diversity, complex platform motion, and realistic trajectories provide a comprehensive resource for future research. Experiments on RealMvMoAT and six public benchmarks show that FUSION achieves state-of-the-art performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑