发表机构
Nanyang Technological University; A*STAR; KAIST; PKU; FDU; University of Catania(南洋理工大学; 新加坡科技研究局; 韩国科学技术院; 北京大学; 复旦大学; 卡塔尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态视频模型在真实第一人称工具使用推理上的不足,提出EgoTools套件,含100小时数据和1000问基准,验证了数据可显著提升模型性能。
AI 中文摘要
从日常活动到专业流程,现实世界中的具身任务要求智能体在物理约束下行动,同时跟踪不断演变的物体和任务状态。工具使用正是此类任务的核心,因为许多日常和专业活动都是通过工具介导的。理解这些活动需要对可供性、手-工具-物体几何关系、程序进展以及对目标物体的因果效应进行推理。然而,尽管当前多模态视频模型在字幕生成和通用视频问答等感知导向的视频任务上表现强劲,但它们在工具中心的具身推理方面仍然能力有限。这一方向的进展一直受到缺乏真实世界第一人称数据和诊断基准的限制。为填补这一空白,我们引入了EgoTools,这是首个面向第一人称工具使用理解的综合套件。它由两个互补的组成部分构成:EgoTools-Data,一个包含100小时工具中心第一人称录制的大型语料库,配有同步音频、密集字幕、推理密集型叙述以及补充3D信息;以及EgoTools-Bench,一个包含1000个问答对、覆盖从感知和几何到程序和因果推理四个轨道的诊断基准。实验结果表明,当前模型仍难以将工具使用与视觉证据相结合:Gemini-3.1-Pro达到66.9%的总体准确率,但在感知与定位轨道上仅为51.7%。除评估外,我们还验证了EgoTools-Data作为训练资源的有效性。在完整的1000问基准上,全监督微调将Qwen3-VL-8B-Instruct从50.0%提升至60.9%,并严格分离源视频。综合这些结果,EgoTools确立了作为真实世界第一人称工具使用理解的训练与诊断评估统一资源的地位。
英文摘要
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
Comments32 pages, 7 figures. Project page: https://ropedia.github.io/egotools