arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ID-VTG:基于图像消歧的视频时间定位

ID-VTG: Image-Disambiguated Video Temporal Grounding

Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu

arXiv 2608.20127首次发表:更新:

发表机构

Wangxuan Institute of Computer Technology, Peking University; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学王选计算机研究所; 北京大学通用人工智能国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频时间定位中视觉相似实体的消歧难题,本文提出ID-VTG任务与VGD-Agg框架,构建两个基准数据集,实现了该任务的当前最优性能。

AI 中文摘要

视频时间定位(Video Temporal Grounding,VTG)面临重大挑战:当自然语言查询需区分涉及视觉相似实体的多个事件时,仅依赖难以用文字准确描述的细粒度视觉属性会导致困难。为解决此问题,本文提出ID-VTG(Image-Disambiguated Video Temporal Grounding)任务,该任务利用结合参考图像与文本描述的多模态查询,精准定位特定实例执行描述动作的片段。为推动研究,本文构建两个基准:IDVTG-Gym,聚焦运动员身着相似制服的细粒度、按顺序组合的体操动作;IDVTG-InternVid,为开放世界数据集,包含多样实体(如人类、动物、虚构角色)及大量时间干扰项。方法上,本文提出基于双分支快慢架构的VGD-Agg(Visually-Guided Disambiguation Aggregation)框架:快分支高效生成初步事件提议,慢分支执行视频帧与参考图像的细粒度帧级匹配。本文通过两个可学习令牌增强判别性:Compare Token,代表难负样本以探测查询图像所指目标实例的存在;Depress Value,代表与文本无关的事件。被Compare Token判定为缺乏目标实例的提议会被推向Depress Value,从而通过文本查询缓解消歧难度。大量实验验证了本文方法,其在提出的基准上取得了当前最优(state-of-the-art)结果。代码可在此https URL获取。

英文摘要

Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual attributes that are difficult to describe accurately in words alone. To address this, we introduce Image-Disambiguated Video Temporal Grounding (ID-VTG), a task that leverages multimodal queries combining a reference image and a text description to precisely localize segments where a specific instance performs a described action. To facilitate research, we construct two benchmarks: IDVTG-Gym, focusing on fine-grained, compositionally ordered gymnastics actions with athletes in similar uniforms; and IDVTG-InternVid, an open-world dataset featuring diverse entities (e.g., humans, animals, fictional characters) and significant temporal distractors. Methodologically, we propose the Visually-Guided Disambiguation Aggregation (VGD-Agg) framework based on a dual-branch fast-slow architecture. The fast branch efficiently generates preliminary event proposals, while the slow branch performs fine-grained frame-level matching between video frames and the reference image. We enhance discriminability via two learnable tokens: a Compare Token, which represents hard negatives to probe for the presence of the target instance (as referred to by the query image), and a Depress Value, which represents text-irrelevant events. Proposals that the Compare Token identifies as lacking the target instance are pushed toward the Depress Value, thus easing disambiguation via the text query. Extensive experiments validate our approach, which achieves state-of-the-art results on the proposed benchmarks. Code is available at https://github.com/oceanflowlab/ID-VTG.

CommentsACM-MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑