arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REZE:基于识别的视频时间定位零样本提取方法

REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding

Boyang Li, Chenhui Gou, Jianfei Cai

arXiv 2608.04480首次发表:更新:

AI 中文总结

该研究提出无训练的REZE方法,通过拆分视频为片段并聚合模型置信度,适配多种VTG任务,在多个数据集上优于现有无训练方法,部分指标超全监督SoTA,还能让旧模型性能接近同家族新模型。

AI 中文摘要

视频时间定位(Video Temporal Grounding,VTG)是指识别视频中与给定自然语言查询对应的时间区间的任务。一种常见的零样本策略是让大型视觉-语言模型(Vision-Language Model,VLM)直接生成起始和结束时间戳,因此结果高度依赖模型的设计与训练,不同VLM的定位精度差异显著。为此,我们提出REcognition-based Zero-shot Extraction(REZE,基于识别的零样本提取),这是一种简单的无训练方法:将视频拆分为短片段,让模型输出片段级别的查询置信度分数,再通过确定性算法将得到的分数曲线转换为任务所需的输出。由于时间聚合操作在模型外部执行,REZE可适配不同的任务输出,涵盖单区间、多区间时刻检索以及亮点检测。在QVHighlights数据集上,REZE将已报道的无训练时刻检索最佳mAP从38.23提升至40.32;在亮点检测任务中,其达到44.18的mAP和73.41的HIT@1,在无训练方法中达到新的最优水平,且其HIT@1指标在QVHighlights测试集上优于所有全监督的当前最优模型(SoTA)。我们在来自三个模型家族的七个主干网络上对REZE进行评估,在Charades-STA和QVHighlights数据集上,REZE在所有可对比项中均优于直接时间戳生成方法。进一步观察发现,借助REZE,较早一代的模型可达到其家族中较新模型的原生性能。

英文摘要

Video temporal grounding (VTG) refers to the task of identifying the time interval in a video that corresponds to a given natural-language query. A common zero-shot strategy asks a large vision-language model (VLM) to generate the start and end timestamps directly, so the result depends heavily on the design and training of the model, and grounding accuracy differs widely from one VLM to another. We therefore propose REcognition-based Zero-shot Extraction (REZE), a simple training-free method that splits the video into short clips, asks the model for a clip-level confidence score for the query, and uses a deterministic algorithm to convert the resulting score curve into the output required by the task. Because temporal aggregation is performed outside the model, REZE adapts to different task outputs, from single- and multi-interval moment retrieval to highlight detection. On QVHighlights, REZE improves the best reported training-free moment-retrieval mAP from 38.23 to 40.32, while on highlight detection it reaches 44.18 mAP and 73.41 HIT@1, establishing a new state of the art among training-free methods. Its HIT@1 also outperforms all fully supervised SoTAs on the QVHighlights test split. We evaluate REZE on seven backbones from three model families. On Charades-STA and QVHighlights, it outperforms direct timestamp generation in every available comparison. We further observe that with REZE an earlier-generation model can approach the native performance of a newer model in its family.

Comments18 pages, 7 figures, 13 tables. Appendices included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑