arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31713cs.CVcs.AI

智能体视频理解:综述

Agentic Video Understanding: A Survey

Xinyu Deng, Siwen Luo, Daochang Liu

AI总结:

本综述系统梳理了以视频为主要信息源的智能体系统,通过挑战-设计分类法分析其关键问题,并展望了智能体原生时序建模与视频原生智能体的未来方向。

AI中文摘要:

随着大型语言模型(LLMs)能够处理日益多样化的模态和更长的时序上下文,一个新兴的研究方向正从固定的视频-语言推理转向智能体系统,这些系统主动决定要检查、保留、验证和采取行动的信息。本综述回顾了视频理解智能体:这些系统以视频为主要信息源,通过自适应状态构建和动作选择来解决理解任务。我们首先形式化了一个用于视频理解的智能体循环,然后解决一个核心问题:为什么智能体对视频理解很重要?为了回答这个问题,我们通过一个从挑战到设计的分类法来组织文献,将上下文瓶颈与分层证据记忆联系起来,将证据稀疏性与主动证据获取联系起来,将时序因果关系与状态和过程跟踪联系起来,将多模态歧义与角色专业化协调联系起来。我们进一步回顾了状态空间范式、学习范式、监督信号、基准和评估协议。最后,我们指出了朝向智能体原生时序建模和视频原生智能体的开放方向。项目页面:此 https URL

英文摘要:

As large language models (LLMs) become capable of processing increasingly diverse modalities and longer temporal contexts, an emerging line of work is moving beyond fixed video-language inference toward agentic systems that actively decide what information to inspect, retain, verify, and act upon. This survey reviews video understanding agents: systems that use video as the primary information source and solve understanding tasks through adaptive state construction and action selection. We first formalize an agent loop for video understanding, then address a central question: why do agents matter for video understanding? To answer this, we organize the literature through a challenge-to-design taxonomy, linking context bottlenecks to hierarchical evidence memory, evidence sparsity to active evidence acquisition, temporal causality to state and process tracking, and multimodal ambiguity to role-specialized coordination. We further review state space paradigms, learning paradigms, supervision signals, benchmarks, and evaluation protocols. Finally, we identify open directions toward agentic-native temporal modeling and video-native agents. Project page: https://github.com/DXY0711/Awesome-Agentic-Video-Understanding

补充信息

↑