发表机构
KRAFTON(魁匠团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentVidBench提出多跳视频问答基准,评估MLLM智能体的时空因果推理,提供解题轨迹,实验显示智能体工作流提升性能,并给出强基线。
AI 中文摘要
全面的视频理解对于推动人工智能向物理世界的复杂动态发展至关重要。尽管多模态大语言模型(MLLMs)的最新进展已在视频理解方面展现出卓越能力,但现有基准仍局限于简单的场景级查询或仅需单步推理的全局摘要。现实世界的视频理解涉及更具挑战性的任务,需要多跳多模态推理,而目前严重缺乏能够严格评估这些智能体能力的视频基准。为弥补这一空白,我们引入了AgentVidBench,一个专注于评估MLLM智能体空间、时间和因果推理能力的多跳视频问答基准。除了标准的问答对,AgentVidBench还提供逐步的解题轨迹以支持轨迹评估,从而判断智能体是否明确获取了证明其答案所需的证据。对12个专有和开源MLLM的实验表明,在AgentVidBench上单轮性能仍然有限,而将这些模型集成到最先进的智能体工作流中,通常在准确性和轨迹得分方面都能提升性能。我们进一步提出了一种简单而有效的智能体策略,作为AgentVidBench上的一个竞争性基线,使我们的基准成为未来智能体视频理解研究的全面测试平台。代码和数据集可在以下网址获取:此https链接和此https链接。
英文摘要
Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these agentic capabilities. To bridge this gap, we introduce AgentVidBench, a multi-hop video question answering benchmark focused on evaluating the spatial, temporal, and causal reasoning capabilities of MLLM agents. Beyond standard question-answer pairs, AgentVidBench provides step-by-step solution traces to support trajectory evaluation that assesses whether agents explicitly acquire the evidence needed to justify their answers. Experiments with 12 proprietary and open-source MLLMs show that single-turn performance remains limited on AgentVidBench, while integrating these models into state-of-the-art agentic workflows generally improves performance with respect to both accuracy and trajectory scores. We further present a simple yet effective agentic strategy that serves as a competitive baseline on AgentVidBench, establishing our benchmark as a holistic testbed for future research on agentic video understanding. Code and datasets are available at https://github.com/krafton-ai/agentvidbench and https://huggingface.co/datasets/agentvidbench/agentvidbench.
Comments36 pages, 8 figures. Code: https://github.com/krafton-ai/agentvidbench Dataset: https://huggingface.co/datasets/agentvidbench/agentvidbench