arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越相似性:基础模型作为免训练组合视频检索的高效骨干

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar, Abdelrahman Mohamed Shaker, Rao Anwer

arXiv 2609.10008首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出免训练组合视频检索框架,通过自适应编排冻结基础模型角色,实现可扩展检索与细粒度推理,在Dense-WebVid-CoVR和CoVR-R上达到最先进性能。

AI 中文摘要

组合视频检索(CoVR)在画廊中搜索目标视频,该视频实现对源剪辑的自然语言修改。然而,在画廊规模下,这产生了根本性的张力:紧凑的嵌入能够实现高效、可复用的搜索,但可能错过需要细粒度视频推理的瞬时动作、状态变化和细微约束,而统一应用大型多模态模型则牺牲了可扩展性。为解决这些限制,我们提出冻结的基础模型应承担互补角色,推理深度根据查询难度进行调整。基于这一前提,我们引入\methodname{},一个免训练的\methodexpansion{}框架。具体而言,组合查询嵌入首先搜索可复用的仅视频画廊表示;不确定的查询进行有界重排序和候选扩展;模糊的编辑触发目标描述生成;只有接近领先的候选才进入多模态验证。为支持这些角色,帧选择、空间分辨率和时间线索适应每个阶段。在完整目标画廊评估中,我们的方法在免训练方法中达到最先进性能,在Dense-WebVid-CoVR和CoVR-R上分别获得89.55和93.43的R@1(与最接近的对应方法相比,绝对边际超过+35%和+25%)。这些结果表明,自适应编排基础模型能力可以结合可扩展检索与细粒度推理,无需任务特定训练。源代码和所有相关指南可在该https URL上获取。

英文摘要

Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constraints that demand fine-grained video reasoning, whereas applying large multimodal models uniformly sacrifices scalability. To address these limitations, we propose that frozen foundation models should instead occupy complementary roles, with inference depth adapted to query difficulty. Based on this premise, we introduce \methodname{}, a framework for training-free \methodexpansion{}. Specifically, a composed-query embedding first searches reusable video-only gallery representations; uncertain queries undergo bounded reranking and candidate expansion; ambiguous edits trigger target-description generation; and only close leading candidates reach multimodal verification. To support these roles, frame selection, spatial resolution, and time cues are adapted to each stage. Across complete target-gallery evaluations, our method reaches state-of-the-art performance among training-free approaches, with 89.55 and 93.43 R@1 on Dense-WebVid-CoVR and CoVR-R, respectively (with more than +35\% and +25\% absolute margins to the closest counterpart). These results show that adaptively orchestrating foundation-model capabilities can combine scalable retrieval with fine-grained reasoning without task-specific training. The source code and all relevant guidelines are available on https://github.com/demidovd98/CoVRAGE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑