RACER:反思式智能体耦合查询解释与基于工具检索的长视频理解帧选择
RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding
浏览论文内容
中文总结 AI 辅助
针对长视频帧选择中的查询理解与解释-选择鸿沟,提出无需训练的反思式智能体框架RACER,通过轻量级Vid-LLM解释查询和嵌入模型检索证据,并迭代反思,提升长视频理解性能。
中文摘要 AI 辅助
视频大语言模型(Vid-LLMs)通过推理选定的帧,在多种视频语言任务中表现出色。然而,长视频的帧选择仍然具有挑战性,因为它需要在给定复杂查询的情况下,从大型候选池中检索分布在各个片段中的相关帧。本文从任务分解的角度研究长视频帧选择的主流方法,识别出两个关键挑战:基于相似度方法中的查询理解鸿沟(Query Comprehension Gap)和基于判断方法中的解释-选择鸿沟(Interpretation--Selection Gap)。为解决这些问题,我们提出RACER,一种无需训练的反思式智能体框架,它将长视频帧选择分解为由轻量级Vid-LLM驱动的查询解释和由作为检索工具的嵌入模型支持的证据定位。具体而言,Vid-LLM仅负责将复杂查询重构为子查询,使隐含的信息需求显式化,从而缓解查询理解鸿沟。同时,检索工具利用这些子查询来定位相关证据,减轻Vid-LLM直接进行帧选择的负担,从而解决解释-选择鸿沟。最后,检索到的帧被反馈给Vid-LLM以进行子查询细化,形成反思循环,迭代改进查询解释和帧选择。跨多个基准的实验表明,RACER持续提升长视频理解性能。值得注意的是,RACER即使使用能力有限的组件也能实现有效的帧选择,这表明智能体集成能使这些组件增强能力更强的Vid-LLM。
英文摘要
Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.
发表机构
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。