arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08850cs.CV

DSE-VTG:用于免训练视频时序定位的双侧增强

DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding

  • The University of Queensland(昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhuo Cao, Bingqing Zhang, Sen Wang, Xue Li

AI总结:

提出DSE-VTG,一种免训练的视频时序定位双侧增强框架,通过多尺度相似性融合和查询级测试时自适应,在多个基准上达到最先进性能。

AI中文摘要:

文本引导的视频时序定位(VTG)旨在根据文本查询在未修剪的视频中定位相关片段,然而收集密集的时间标注和训练特定任务的模型在分布偏移下仍然代价高昂且脆弱。最近的免训练VTG方法通过直接匹配预训练的视觉-语言表示来缓解这一问题,但它们仍面临两个根本性的信息瓶颈:逐帧的视觉编码忽视了时间动态,而固定的查询嵌入无法解决查询歧义。为解决这些问题,我们提出了DSE-VTG,一个双侧增强框架,在无需任何任务特定训练的情况下同时解决这两个问题。在视觉侧,多尺度相似性融合(MSF)将帧级和片段级相似性结合成一个统一的、具有时间感知的相似性剖面。在文本侧,查询级测试时自适应(Q-TTA)优化一个轻量级的加性偏移,在测试时将查询嵌入适应到视频,无需微调骨干网络或调用外部大型语言模型。在三个标准基准和两个OOD基准上的大量实验表明,DSE-VTG在免训练方法中达到了最先进的性能。在Charades-STA上,它比最强的先前免训练方法将mIoU提高了5.61个百分点。在分布偏移下,DSE-VTG在Charades-CG Novel-Word上达到了50.86的mIoU,超过了最强的监督基线2.76个mIoU。我们的代码将在录用后发布。

英文摘要:

Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based on text queries, yet collecting dense temporal annotations and training task-specific models remain costly and brittle under distribution shift. Recent training-free VTG approaches mitigate this issue by directly matching pretrained vision-language representations, but they still face two fundamental information bottlenecks: frame-wise visual encoding overlooks temporal dynamics, while fixed query embeddings cannot resolve query ambiguity. To address these issues, we propose DSE-VTG, a \underline{D}ual-\underline{S}ide \underline{E}nhancement framework that addresses both without any task-specific training. On the visual side, Multi-scale Similarity Fusion (MSF) combines frame- and clip-level similarities into a unified, temporally aware similarity profile. On the textual side, Query-level Test-Time Adaptation (Q-TTA) optimizes a lightweight additive offset to adapt the query embedding to the video at test time, without finetuning the backbone or calling external large language models. Extensive experiments on three standard and two OOD benchmarks show that DSE-VTG achieves state-of-the-art performance among training-free methods. On Charades-STA, it improves mIoU over the strongest prior training-free method by 5.61 points. Under distribution shift, DSE-VTG reaches 50.86 mIoU on Charades-CG Novel-Word, surpassing the strongest supervised baseline by 2.76 mIoU. Our code will be released upon acceptance.

补充信息

↑