发表机构
Nanyang Technological University; Mohamed bin Zayed University of Artificial Intelligence; University of North Carolina at Chapel Hill(南洋理工大学; 穆罕默德·本·扎耶德人工智能大学; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Ego2Act提出首个目标导向第一人称视频生成基准,含2640个真实任务视频,并配套无参考评估器Ego2ActJudge,揭示现有模型在多步物理推理和操作细节上的不足。
AI 中文摘要
视频生成模型正越来越多地被探索作为具身规划和学习的世界模拟器。为了有效实现这一点,这些模型不仅必须生成视觉上吸引人的帧,还必须预测在执行目标导向动作时环境如何动态演变。虽然评估这些能力至关重要,但现有基准主要关注单个短动作或逐步指令,这使得多步骤物理推理未被充分探索,尤其是在需要规划以通过执行多个真实世界操作来实现高层目标的第一人称视频生成中。我们引入了Ego2Act,一个目标导向基准,包含来自110个日常场景真实任务的2640个视频,具有不同的物体杂乱程度和多步骤复杂性。给定初始场景图像和高级目标,Ego2Act评估视频生成模型能否生成手部操作物体以执行任务的逼真第一人称视频。为了支持可扩展评估,我们还引入了Ego2ActJudge,一个无参考评估流程,与相关基线相比,在任务完成和物理合理性评估上与人类共识实现了更好的一致性。我们的发现表明,模型生成的模拟经常跳过或部分执行步骤,导致后续步骤缺少依赖状态,从而未能实现目标。此外,模型在细粒度物理动态方面持续失败,特别是在复杂物体操作和持久世界建模中。我们希望Ego2Act为推进视频模型向物理上合理、目标导向的模拟发展提供一个严格的测试平台。
英文摘要
Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.
CommentsPreprint. 51 pages, 19 figures, 23 tables. Code, dataset and project website linked in the paper