arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EgoErrorVQA:通过程序错误评估自我主体人工智能的以自我为中心的理解能力

EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI

Junlong Li, Junxi Li, Jianjun Gao, Chen Cai, Lap-Pui Chau, Yi Wang

arXiv 2608.24134首次发表:更新:

发表机构

The Hong Kong Polytechnic University; Nanyang Technological University(香港理工大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视觉智能体和VLMs忽略以自我为中心视角评估程序理解能力的问题,本文提出EgoErrorVQA任务,开发基于A2A协议的评估智能体,引入Ego-ADR框架提升模型对程序错误的理解,取得良好效果。

AI 中文摘要

我们的大部分日常活动都是程序性的,由一系列相互依赖的步骤组成。然而,现有的视觉智能体和视觉语言模型(VLMs)基准忽略了从以自我为中心的视觉视角评估其程序理解能力,尤其是检测程序错误这一对日常辅助至关重要的能力。为填补这一空白,本文首次提出EgoErrorVQA任务,用于以自我为中心的程序理解,并对程序错误进行显式建模。此外,我们基于Agent2Agent(A2A)协议开发了一款用户友好的评估智能体,通过基于VQA的交互实现对视觉智能体的严格、标准化评估。我们使用开放式和多项选择题对一系列模型进行评估,结果显示这些模型在处理程序错误及错误类型方面存在持续的弱点。此外,我们引入Ego-ADR,一种自适应解耦推理框架,该框架将复杂的程序推理解耦,以增强模型对程序错误的理解。在可比设置下,该框架在选定基准上实现了稳定的性能提升,并在多项指标上达到了最先进的结果。代码:this https URL

英文摘要

The majority of our everyday activities are procedural and consist of sequences of interdependent steps. However, existing benchmarks for Visual Agents and Visual Language Models (VLMs) overlook the evaluation of their procedural comprehension ability from an egocentric visual perspective, particularly for detecting procedural errors, a critical capability for everyday assistance. To bridge this gap, the EgoErrorVQA task is firstly proposed for egocentric procedural comprehension with explicit procedural errors modeling. Besides, we develop a user-friendly evaluator agent based on the Agent2Agent (A2A) protocol, enabling rigorous and standardized evaluation of visual agents through VQA-based interaction. A range of models are evaluated using both open-ended and multiple-choice questions, revealing persistent weaknesses in handling procedural errors and error types. Moreover, we introduce Ego-ADR, an Adaptive Decoupled Reasoning framework that decouples complex procedural reasoning to enhance models' understanding of procedural errors. It achieves consistent performance gains over the selected baselines and attains state-of-the-art results on several metrics under comparable settings. Code: https://github.com/z1oong/EgoErrorVQA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑