arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VEDJE:用于可扩展视频-文本检索的视频高效判别式联合编码器

VEDJE: Video-Efficient Discriminative Joint Encoder for Scalable Video-Text Retrieval

Shahaf Wagner, Gabriele Serussi, Dan Ben Ami, Tomer Galanti, Chaim Baskin

arXiv 2610.11850首次发表:更新:

发表机构

Ben-Gurion University of the Negev; Texas A&M University; Decart AI(内盖夫本-古里安大学; 德克萨斯农工大学; 德卡特人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出VEDJE模型,通过压缩视频帧特征并结合特征变化预测辅助训练,在多个视频-文本检索数据集上提升R@1指标,缩小缓存规模后仍保持较高召回率,实现高效可扩展的视频检索。

AI 中文摘要

查找合适的视频通常需要区分包含不同事件的相似场景。联合匹配可提升检索效果,但为每个查询处理丰富的视频表示成本高昂。VEDJE在采样帧内压缩特征,同时将这些特征的表示保留在可复用缓存中。特征变化预测提供辅助训练信号,可提升压缩缓存的检索效果,且不会在查询时增加额外工作量。在MSR-VTT、MSVD、DiDeMo和ActivityNet数据集上,VEDJE在两种检索方向上均优于匹配的第一阶段检索器,提升了R@1指标。在MSR-VTT数据集上,当使用微调后的VideoCLIP-XL作为第一阶段时,VEDJE达到了59.8的文本到视频R@1。在VideoPrism配置下,将每个视频的缓存缩小四倍至12 KiB时,文本到视频的召回率仍保持在0.2个百分点以内。这些结果表明,准确的视频搜索可基于紧凑证据运行,该证据经一次编码后,可在新查询到达时复用。

英文摘要

Finding the right video often requires distinguishing similar scenes in which different events occur. Joint matching improves retrieval, but processing rich video representations for each query is costly. VEDJE compresses features within sampled frames while keeping their representations separate in a reusable cache. Feature-change prediction supplies an auxiliary training signal that improves retrieval from the compressed cache without adding work at query time. On MSR-VTT, MSVD, DiDeMo, and ActivityNet, VEDJE improves R@1 over matched first-stage retrievers in both retrieval directions. On MSR-VTT, it reaches 59.8 text-to-video R@1 with a fine-tuned VideoCLIP-XL first stage. In the VideoPrism configuration, shrinking the per-video cache fourfold to 12 KiB preserves text-to-video recall within 0.2 points. These results show that accurate video search can operate on compact evidence, encoded once and reused as new queries arrive.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑