VideoGAIA:面向通用人工智能助手的智能体视频理解基准
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
浏览论文内容
中文总结 AI 辅助
VideoGAIA是面向通用AI助手的智能体视频理解基准,构建多轮工具增强交互任务,现有MLLMs在其上准确率不足60%,可评估下一代模型并推动相关转型。
中文摘要 AI 辅助
视频理解是评估多模态大语言模型(MLLMs)能力的基础任务。然而,现有领先模型在Video-MME排行榜上已达到约90%的准确率,表明传统的单轮视频理解任务正逐渐饱和,不足以评估先进MLLMs的智能水平。为此,我们推出VideoGAIA——一个面向通用人工智能(AI)助手的智能体视频理解基准。该基准突破了单轮视频问答的局限,将视频理解构建为多轮、工具增强的交互过程,要求模型迭代感知视频、调用外部工具、收集补充信息并整合多轮多模态证据。VideoGAIA包含271项由模型与人类共同设计的任务,涵盖多样且复杂的真实场景;每个视频-问题-答案实例均由三名人类专家独立验证,以确保正确性与难度适配。包括GPT-5.5、Kimi-K3在内的所有参与评估的MLLMs在VideoGAIA上的准确率均低于60%,凸显其作为高质量、时效性基准在评估下一代MLLMs方面的价值。我们希望VideoGAIA能推动从传统视频理解向智能体视频理解的转型。
英文摘要
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Ant Group(蚂蚁集团)
- Tsinghua University(清华大学)
- Tongji University(同济大学)
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。