发表机构
China University of Mining and Technology; Mine Digitization Engineering Research Center of the Ministry of Education; University of Auckland(中国矿业大学; 矿山数字化教育部工程研究中心; 奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ESMTrack,一种无需模态误导的端到端自监督RGB-T跟踪框架,利用接地与时间三元组损失及模态解耦机制,在五个基准上达到先进性能。
AI 中文摘要
RGB-T目标跟踪利用可见光和热红外模态的互补特性,以提高在不利条件下的鲁棒性。现有的监督方法通常依赖昂贵的模态对齐边界框标注,而大多数自监督方法遵循两阶段伪标签范式,使得跟踪器训练对伪标签质量敏感,并阻碍了端到端的联合优化。本文提出ESMTrack,一种完全端到端的自监督RGB-T跟踪框架,无需离线伪标签生成或密集帧级边界框标注。仅利用视觉跟踪中标准的初始帧标注,ESMTrack通过两个互补目标学习判别性和时间一致性表示:在标注的初始帧上的接地三元组损失,以及在未标记的搜索帧上的跨帧时间三元组损失,并通过前向-后向一致性选择可靠样本。为解决模态主导偏差,ESMTrack采用三分支架构,包括一个融合分支和两个分别用于RGB和热输入的单模态分支。我们通过测量融合分支与单模态分支之间的响应差异,使用平均峰值相关能量来量化模态贡献。由此产生的可靠性估计指导训练时的模态解耦机制,抑制主导模态捷径,并自适应加权跨模态对比学习以实现任务级对齐。在五个RGB-T跟踪基准上的大量实验表明,ESMTrack实现了具有竞争力的最先进性能、强大的跨数据集泛化能力和实时推理速度。源代码可在以下网址获取:此https URL。
英文摘要
RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at https://github.com/LiShenglana/ESMTrack.