arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨模态无人机目标跟踪:状态感知表示学习与统一基准

Cross-Modal UAV Object Tracking: State-Aware Representation Learning and A Unified Benchmark

Yun Xiao, Zhihong Hong, Jiandong Jin, Chenglong Li, Jin Tang, Amir Hussain

arXiv 2607.18768首次发表:更新:

发表机构

Anhui University; Edinburgh Napier University(安徽大学; 爱丁堡龙比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无人机跨模态目标跟踪中模态切换致外观和位置突变的问题,提出状态感知表示学习方法SARLA,含相关模块及损失函数,建立CM-UOT基准,实验证明其性能优于多种方法,推动了该领域发展。

AI 中文摘要

无人机目标跟踪成为热门研究领域且应用广泛。现代无人机常配备可见光和热红外传感器,但因通信带宽等限制,当前系统常激活一种模态并切换,导致外观变化和空间偏移,给跟踪算法带来挑战。为此提出状态感知表示学习方法SARLA,含模态状态感知表示模块和空间状态感知表示模块,还设计空间偏移预测损失。建立CM-UOT基准,含1079个跨模态序列。实验表明SARLA性能优于20种优秀跟踪方法。

英文摘要

Unmanned Aerial Vehicle (UAV) object tracking has emerged as a popular research field with broad practical applications. Modern UAVs are increasingly equipped with both visible light and thermal infrared sensors. However, due to constraints in communication bandwidth, computational resources and power consumption, current systems often activate one modality and switch between modalities to maintain robust tracking in complex scenarios. Such modality switch inevitably leads to significant appearance change and sudden spatial shift, posing great challenges for existing tracking algorithms. To handle this problem, we propose a novel State-Aware Representation Learning Approach called SARLA, which perceives the inconsistent modality states of current frame with template and last frame in the target representations to adapt to the sudden changes in both appearance and position, for robust cross-modal object tracking. In particular, we propose the Modality State Aware Representation Module (MSARM) and Spatial State Aware Representation Module (SSARM). MSARM guides the model to learn appearance correlation, bridging the modality gap, while SSARM models cross-frame spatial correlation to mitigate sudden spatial shift impacts. In addition, we design a spatial shift prediction loss to further handle the effects of spatial variation caused by modality switch. To promote the development of this research field, we establish a large-scale video benchmark called CM-UOT, which consists of 1079 cross-modal sequences with an average video length greater than 621 frames and encompasses over 671K frames in total. Extensive experiments on CM-UOT dataset demonstrate the superior performance of the proposed SARLA against 20 excellent tracking methods. The source code, datasets, and evaluation protocols associated with this work are publicly available at: https://github.com/hongsmile365/sarla-.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑