AI 中文总结
本文研究现代视觉语言模型智能体能否仅凭初始搜索意图描述,自主操作交互式视频检索系统,并在多种设置下达到与专家系统竞争的性能。
AI 中文摘要
搜索大型视频集通常是一个交互式过程,用户在其中扮演两个角色。首先,他们持有搜索意图:即决定他们寻找什么内容以及为何寻找的潜在目标。其次,用户必须通过迭代搜索循环来操作化这一意图。用户将意图转化为查询,浏览检索到的候选结果,并根据结果优化查询。在本文中,我们研究了现代视觉语言模型(VLM)和智能体方法以交互式且完全自主的方式达成搜索目标的能力。具体而言,我们研究了所提供的搜索目标的初始描述是否足以通过智能体系统解决传统上交互式的搜索任务。鉴于所涉及的VLM事先并不了解整个大型视频数据集,关键挑战在于如何有效结合现有的交互式视频搜索系统和控制该系统的智能VLM智能体。搜索系统提供索引和高效查询,而基于VLM的智能体分析排名靠前的条目并决定下一步行动。我们的结果表明,现代智能体能够自主操作交互式视频检索系统,从初始意图描述中解决许多搜索任务,在多种设置下达到与强大的历史专家操作系统竞争的性能。
英文摘要
Searching large video collections is typically an interactive process in which users play two roles. First, they hold the search intent: the underlying goal that determines what content they seek and why. Second, users must operationalize this intent through an iterative search loop. Users translate their intent into queries, browse the retrieved candidates, and refine their queries based on the results. In this paper, we investigate the capabilities of modern Vision Language Models (VLM) and agentic approaches to reach search goals interactively and fully autonomously. Specifically, we study whether a provided initial specification of a search goal might be sufficient to solve traditionally interactive search tasks with an agentic system. Provided that the involved VLMs are not aware of the whole large video dataset in advance, the key challenge lies in the effective combination of an existing interactive video search system and a smart VLM agent controlling the system. While the search system provides indexing and efficient querying, the VLM-based agents analyze top-ranked items and make decisions about next actions. Our results show that modern agents can autonomously operate interactive video retrieval systems to solve many search tasks from an initial intent description, achieving performance competitive with strong historical expert-operated systems in several settings.
CommentsSubmitted to the International Conference on Multimedia Modeling