用于指称单目标跟踪的高效语言到视觉特征注入
Efficient Language-to-Vision Feature Injection for Referring Single-Object Tracking
- Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间应用工程与技术中心)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出LVTrack纯Transformer框架,通过模式条件门控特征注入器等设计缓解语义漂移,利用冻结的视觉-语言预训练模型降低训练成本,在指称单目标跟踪任务的标准基准上取得优异性能。
AI中文摘要:
指称单目标跟踪通过联合利用语义线索与视觉模板,实现基于语言的目标初始化及后续跟踪。核心难点在于需在不同阶段差异化使用语言:语言对目标定位不可或缺,但过度强调会在跟踪过程中引发语义漂移。同时,现有方法通常需要昂贵的视觉-语言对齐训练。本文提出LVTrack,一种纯Transformer框架,引入模式条件门控特征注入器以自适应调节文本引导、缓解语义漂移;结合针对性适配,直接利用冻结的视觉-语言预训练模型,大幅降低训练成本并保留强大的语言理解能力。为进一步提升时间定位性能,LVTrack融合混合相对-绝对位置编码与轻量记忆机制,并采用高斯平滑KL损失优化自回归框预测。在标准基准上开展的大量实验表明,LVTrack取得了优异性能。
英文摘要:
Referring single-object tracking enables language-grounded target initialization and subsequent tracking by jointly leveraging semantic cues and visual templates. The core difficulty is to use language differently across stages: it is indispensable for grounding but can induce semantic drift during tracking when overemphasized. Meanwhile, current methods often require costly vision-language alignment training. We present LVTrack, a pure transformer framework that introduces a mode-conditioned Gated Feature Injector to adaptively regulate textual guidance and alleviate semantic drift. Together with targeted adaptations, it directly harnesses a frozen vision-language pretrained model, greatly reducing training cost and preserving strong language understanding. To further improve temporal localization, LVTrack integrates hybrid relative-absolute positional encodings with a lightweight memory mechanism and optimizes autoregressive box prediction using a Gaussian-smoothed KL loss. Extensive experiments on standard benchmarks demonstrate that LVTrack achieves strong performance.