arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于模态缺失的RGBT跟踪的时空条件去噪Transformer

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Andong Lu, Ziyi Zha, Jiandong Jin, Shihao Li, Chenglong Li, Jin Tang, Bin Luo

arXiv 2607.24701首次发表:更新:

发表机构

School of Computer Science and Technology, Anhui University; School of Artificial Intelligence, Anhui University(安徽大学计算机科学与技术学院; 安徽大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RGBT跟踪中模态缺失致性能下降问题,提出时空条件去噪Transformer(SCDT)。它整合时空线索,通过利用不同时间线索及噪声调制适应机制,在统一框架下实现缺失模态信息重建等,实验证明其性能优于现有方法。

AI 中文摘要

在RGBT跟踪中,模态缺失往往会导致多模态特征表示不完整和不稳定,从而严重降低性能。现有方法通常试图从可用模态中恢复缺失模态,但在具有挑战性的场景中生成的数据质量可能不尽人意。此外,当前方法在处理缺失和完整数据时灵活性有限。为克服这些限制,我们提出了时空条件去噪Transformer(SCDT),它在统一框架中整合空间线索和时间上下文,以自适应地对缺失模态进行信息重建和对弱模态进行特征增强,用于稳健的模态缺失RGBT跟踪。具体而言,SCDT利用近期历史帧的短期时间线索捕捉细粒度时间相关性,利用编码模态演变的长期时间线索捕捉全局上下文。通过联合利用长期和短期时间上下文作为条件,SCDT逐步引导可用模态的噪声特征学习可靠且时间一致的多模态表示。此外,SCDT引入了噪声调制适应机制,可根据模态可用性动态调整其行为,使单个框架能够在模态缺失和完整场景下统一特征学习,而无需更改架构或参数。在三个公共基准数据集上的大量实验表明,我们的方法始终优于现有方法。代码可在此处获取。

英文摘要

Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition, current approaches exhibit limited flexibility in processing both missing and complete data. To overcome these limitations, we propose a Spatio-temporal Conditional Denoising Transformer (SCDT), which integrates the spatial cues and the temporal context to adaptively perform information reconstruction of missing modalities and feature enhancement of weak modalities in a unified framework, for robust modality-missing RGBT tracking. In particular, SCDT leverages the short-term temporal cues from recent historical frames to capture the fine-grained temporal correlations and the long-term temporal cues encoding modality evolution to capture the global context. By jointly exploiting long short-term temporal contexts as the conditions, SCDT progressively guides noisy features of available modalities to learn reliable and temporally consistent multimodal representations. Furthermore, SCDT introduces a noisemodulated adaptation mechanism that dynamically adjusts its behavior according to the modal availability, enabling a single framework to unify feature learning under both modality-missing and complete scenarios without changing the architecture or parameters. Extensive experiments on three public benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods. The code is available here.

CommentsAccepted by CVPR2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑