arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAM:面向实体中心视频的连续提取与自适应查询问答

CAM: Question Answering on Entity-Centric Videos with Continuous Extraction and Adaptive Querying

Yizhou Tian, Zizhe Chen, Shiyuan Deng, Garry Yang, Zijie Dai, Luohao Pan, Hao Lin, Peiqi Yin, Xiao Yan, James Cheng

arXiv 2609.06504首次发表:更新:

发表机构

The Chinese University of Hong Kong; Knowin AI; Institute for Math and AI, Wuhan University(香港中文大学; 知因智芯(Knowin AI); 武汉大学数学与人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CAM通过知识图谱连续提取高层语义,并利用规划器-执行器-验证器自适应组合多种搜索方法,提升长视频实体中心问答的细粒度检索,在三个基准上超越SOTA,准确率最高提升23个百分点。

AI 中文摘要

记忆通过提取和检索事实,将多模态大语言模型(MLLMs)有限的上下文窗口内的信息进行整合,从而促进对长视频的问答。现有解决方案通常从固定长度的视频片段中提取独立的记忆条目,因此无法捕获需要在较长时间段内进行总结的高层语义,例如角色特征和关系。此外,它们仅依赖基于相似性的检索,可能无法检索到回答问题所需的细粒度细节。为解决这些问题,我们提出了CAM,其特点是针对高层语义进行连续提取,并针对细粒度细节进行自适应查询。具体而言,CAM将从视频片段中提取的实体和关系存储在知识图谱中。为了捕获每个实体或关系的高层语义,一旦目标实体或关系的局部子图达到预定义大小,CAM便对该子图进行总结。为了检索回答问题所需的细粒度细节,CAM支持多种搜索方法,包括知识图谱遍历、视频重看和音频收听。它利用规划器-执行器-验证器流水线,根据问题意图自适应地组合这些搜索方法。在三个基准上的评估表明,CAM优于最先进的基线,并将其准确率提高了最多23个百分点。代码可在以下网址获取:此https URL。

英文摘要

Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions typically extract independent memory entries from fixed-length video clips and thus cannot capture high-level semantics that need to be summarized over extended time periods, such as character traits and relations. Moreover, they rely solely on similarity-based retrieval and may fail to retrieve the fine-grained details required for question answering. To tackle these problems, we propose CAM, featuring continuous extraction for high-level semantics and adaptive querying for fine-grained details. In particular, CAM stores the entities and relations extracted from video clips in a knowledge graph. To capture the high-level semantics of each entity or relation, CAM summarizes the local subgraph of the target entity or relation once the subgraph reaches a predefined size. To retrieve the fine-grained details required for question answering, CAM supports multiple search methods, including knowledge graph traversal, video re-watching, and audio listening. It utilizes a planner-executor-verifier pipeline to adaptively compose these search methods according to question intent. Evaluations on three benchmarks show that CAM outperforms SOTA baselines and improves their accuracy by up to 23 percentage points. Code is available at https://github.com/Jake-Tian/CAM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑