YesTrack:基于多模态大语言模型(MLLM)是/否验证的指称多目标跟踪
YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification
- University of Electronic Science and Technology of China(电子科技大学)
- Shenzhen Institute for Advanced Study, UESTC(电子科技大学深圳高等研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对现有RMOT方法浪费MLLM能力的局限,提出YesTrack将指称任务转为MLLM是/否验证的判别式任务,引入TCP、TRP约束,在Refer-KITTI等数据集上实现SOTA性能且效率优异。
AI中文摘要:
指称多目标跟踪(RMOT)旨在跟踪视频中所有符合给定语言表达式的实例。尽管近期有研究将多模态大语言模型(MLLM)融入以提升泛化能力,但现有方法大多仅将其用作字幕生成器,需外部模块完成最终决策,这不仅增加了额外延迟,还严重浪费了MLLM固有的视觉-语言对齐能力。为解决这些局限,我们提出YesTrack,一种新型两阶段RMOT方法,将指称任务重新定义为判别式任务,无需显式文本生成,直接利用MLLM完成是/否验证。为进一步提升该MLLM验证的可靠性与效率,我们引入两个轻量型时间一致性约束:时间置信度先验(TCP)与时间指称传播(TRP)。我们还通过提出YesTrack-MOT(针对通用多目标跟踪(MOT)的直接且高效的实例)验证了该判别式范式的通用性。在Refer-KITTI与Refer-KITTI-V2上的实验表明,即便采用最小规模的Qwen3-VL实现,YesTrack仍显著优于现有SOTA方法,且保持高推理效率。代码已在该httpsURL开源。
英文摘要:
Referring multi-object tracking (RMOT) aims to track every instance in a video that matches a given language expression. Despite the recent integration of multimodal large language models (MLLMs) to enhance generalization, existing methods predominantly relegate them to the role of caption generators, necessitating external modules for final decision-making. This paradigm not only introduces extra latency but also severely underutilizes the inherent vision-language alignment capabilities of MLLMs. To address these limitations, we propose YesTrack, a novel two-stage RMOT method that reformulates referring as a discriminative task, directly leveraging MLLMs for Yes/No verification without explicit text generation. To further enhance the reliability and efficiency of this MLLM-based verification, we introduce two lightweight temporal consistency constraints: Temporal Confidence Prior (TCP) and Temporal Reference Propagation (TRP). We further validate the generality of this discriminative paradigm by proposing YesTrack-MOT, a straightforward yet highly effective instantiation for generic multi-object tracking (MOT). Experiments on Refer-KITTI and Refer-KITTI-V2 show that YesTrack significantly outperforms existing state-of-the-art methods while maintaining high efficiency, even when implemented with the smallest variant of Qwen3-VL. Code is released at https://github.com/ggbondrighthere24/YesTrack.