arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29126cs.CV

用于指称单目标跟踪的高效语言到视觉特征注入

Efficient Language-to-Vision Feature Injection for Referring Single-Object Tracking

  • Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间应用工程与技术中心)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

Han Wang, Yuxuan Liu, Yuhan Sun, Jian Yang, Xiaotong Xu, Yixuan Lv, Zhuang Zhou, Shengyang Li

AI总结:

本文提出LVTrack纯Transformer框架,通过模式条件门控特征注入器等设计缓解语义漂移,利用冻结的视觉-语言预训练模型降低训练成本,在指称单目标跟踪任务的标准基准上取得优异性能。

AI中文摘要:

指称单目标跟踪通过联合利用语义线索与视觉模板,实现基于语言的目标初始化及后续跟踪。核心难点在于需在不同阶段差异化使用语言:语言对目标定位不可或缺,但过度强调会在跟踪过程中引发语义漂移。同时,现有方法通常需要昂贵的视觉-语言对齐训练。本文提出LVTrack,一种纯Transformer框架,引入模式条件门控特征注入器以自适应调节文本引导、缓解语义漂移;结合针对性适配,直接利用冻结的视觉-语言预训练模型,大幅降低训练成本并保留强大的语言理解能力。为进一步提升时间定位性能,LVTrack融合混合相对-绝对位置编码与轻量记忆机制,并采用高斯平滑KL损失优化自回归框预测。在标准基准上开展的大量实验表明,LVTrack取得了优异性能。

英文摘要:

Referring single-object tracking enables language-grounded target initialization and subsequent tracking by jointly leveraging semantic cues and visual templates. The core difficulty is to use language differently across stages: it is indispensable for grounding but can induce semantic drift during tracking when overemphasized. Meanwhile, current methods often require costly vision-language alignment training. We present LVTrack, a pure transformer framework that introduces a mode-conditioned Gated Feature Injector to adaptively regulate textual guidance and alleviate semantic drift. Together with targeted adaptations, it directly harnesses a frozen vision-language pretrained model, greatly reducing training cost and preserving strong language understanding. To further improve temporal localization, LVTrack integrates hybrid relative-absolute positional encodings with a lightweight memory mechanism and optimizes autoregressive box prediction using a Gaussian-smoothed KL loss. Extensive experiments on standard benchmarks demonstrate that LVTrack achieves strong performance.

↑