发表机构
Army Engineering University of PLA; Dalian University of Technology; National University of Defense Technology(中国人民解放军陆军工程大学; 大连理工大学; 国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多模态目标跟踪器适配性与泛化性差的问题,提出AnyTrack框架,通过MIM与CUM模块实现任意模态的统一跟踪,扩展基准后实验验证其性能先进、灵活有效。
AI 中文摘要
视觉目标跟踪旨在连续定位序列帧中的特定目标,已从单模态方法发展至多模态方法。然而,现有多模态跟踪器通常为固定模态组合设计,需为不同输入单独构建模型,这导致其对缺失或不完美模态的适应性差,泛化能力有限。为解决这些问题,我们提出名为AnyTrack的新型统一框架,用于任意模态的目标跟踪。具体而言,我们设计了模态感知交互模块(Modality-aware Interaction Module, MIM)以促进不同模态间的动态交互,该模块弥合模态差异并聚合时序线索,在跨模态交互期间保持时空一致性。此外,我们引入上下文理解模块(Context Understanding Module, CUM),通过全局-局部提示建立视觉特征与目标位置间的空间对应关系,该模块采用目标感知上下文建模以增强前景-背景区分,实现精确定位。最后,为支持不同模态下的训练与评估,我们通过加入灰度图像、语言描述和音频片段扩展了现有的多模态目标跟踪基准。在完整模态和缺失模态设置下的大量实验表明,我们的AnyTrack实现了最先进的性能,验证了其有效性和灵活性。源代码可在该https URL获取。
英文摘要
Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single-modal methods to multi-modal ones. However, existing multi-modal trackers are typically designed for fixed modality combinations, requiring separate models for different inputs. This leads to a poor adaptability to missing or imperfect modalities, and limited generalization. To address these issues, we propose a novel unified framework called AnyTrack for object tracking with any modalities. Specifically, we design a Modality-aware Interaction Module (MIM) to facilitate dynamic interaction across diverse modalities. This module bridges modality discrepancies and aggregates temporal cues to maintain spatio-temporal consistency during cross-modal interaction. Furthermore, we introduce a Context Understanding Module (CUM) to establish spatial correspondence between visual features and target locations via global-local prompts. This module employs target-aware context modeling to enhance foreground-background discrimination for precise localization. Finally, to support the training and evaluation under diverse modalities, we extend existing multi-modal object tracking benchmarks by incorporating grayscale images, language descriptions, and audio clips. Extensive experiments with both complete and missing modality settings demonstrate that our AnyTrack achieves state-of-the-art performance, validating its effectiveness and flexibility. The source code is available at https://github.com/IdolLab/AnyTrack.
CommentsAccepted by ACM MM2026. More modifications may be performed