发表机构
The Hong Kong Polytechnic University; South China University of Technology(香港理工大学; 华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有自我中心视频理解评估忽视程序理解的问题,引入EgoProceVQA任务,开发EgoProceGen数据生成平台并构建基准。通过评估发现模型不足,进而提出EgoProceAgent框架,设计相关工具库,使其在多任务上获最优性能,为该任务奠定统一基础。
AI 中文摘要
大多数日常活动本质上都是程序性的。然而,现有的自我中心视频理解评估很少涉及程序理解,在广泛使用的多模态语言模型(MLLMs)的视频问答(VQA)范式下,很大程度上忽略了复杂的关键步骤级推理。为填补这一空白,我们引入了自我中心程序理解VQA任务(EgoProceVQA),通过六种以关键步骤为中心的问题系统评估当前MLLMs和代理的自我中心程序推理能力。此外,我们开发了EgoProceGen数据生成平台,构建了一个包含3600个问题、四种常见程序场景和31个日常程序任务的基准。评估表明现有模型在程序理解方面仍有很大提升空间。因此,我们进一步提出EgoProceAgent自我技能探索代理框架,设计了通用工具库和标准化子技能库,使其能在无监督下自我探索,在多个任务上取得了开源模型中的最优性能。我们的基准、生成平台和代理框架为EgoProceVQA建立了统一基础。
英文摘要
Most daily activities are inherently procedural. However, existing evaluations for egocentric video understanding seldom address procedural understanding and largely overlook complex key-step-level reasoning under the widely used video question answering (VQA) paradigm for MLLMs. Such capabilities are crucial for building procedural AI assistants deployable on wearable devices. To bridge this gap, we introduce the Egocentric Procedural Understanding VQA task (EgoProceVQA), which systematically evaluates egocentric procedural reasoning abilities of current MLLMs and agents through six types of key-step-centric questions. Furthermore, we develop EgoProceGen, a data generation platform that efficiently constructs QA data tailored to different question types. Based on this platform, we build a benchmark with 3,600 questions, four common procedural scenarios, and 31 everyday procedural tasks. Evaluations on EgoProceVQA show that existing MLLMs and agents still have substantial room for improvement in procedural understanding. Therefore, we further propose EgoProceAgent, a self-skill-exploration agentic framework. We design a generic tool library for procedural understanding and a standardized sub-skill library shared across tools and models, enabling self-exploration without ground-truth supervision. By exploring how to compose and select sub-skills, the agent discovers effective skill strategies for diverse problems, and attains state-of-the-art performance among open-source models on multiple tasks. Together, our benchmark, generation platform, and agentic framework establish a unified foundation for EgoProceVQA. Project page: https://z1oong.github.io/EgoProceVQA/.