REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding
REZE:基于识别的视频时间定位零样本提取方法
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract,abstract_cn);VLM(abstract,abstract_cn);分类 cs.CV
AI总结 该研究提出无训练的REZE方法,通过拆分视频为片段并聚合模型置信度,适配多种VTG任务,在多个数据集上优于现有无训练方法,部分指标超全监督SoTA,还能让旧模型性能接近同家族新模型。
Comments 18 pages, 7 figures, 13 tables. Appendices included