Video2World:从具身视频评测编码智能体进行交互式世界建模
Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos
浏览论文内容
中文总结 AI 辅助
本文提出Video2World基准,将自主视频到模拟任务形式化,评测编码智能体从具身视频构建交互式模拟世界的能力,发现模型性能提升但视觉保真度与任务成功率存在差距。
中文摘要 AI 辅助
从真实世界观察中构建交互式模拟器是扩展具身数据的一种有前景的方法,但当前的流程仍然严重依赖人工环境构建和校准。我们研究前沿基础模型和编码智能体能否端到端地自动化这一过程。我们将自主视频到模拟(autonomous video-to-simulation)形式化为一项软件工程任务,其中智能体观察具身视频,构建相应的模拟环境和机器人行为,并通过执行反馈迭代地优化结果。为了评估这一能力,我们引入了Video2World基准,该基准包含来自189个机器人和人类演示视频的222个重建实例。Video2World从几何保真度、动态保真度和功能正确性三个方面衡量重建世界,捕捉空间感知、物理推理和可执行交互。对9个前沿编码智能体系统的评估显示,从Claude Opus 5开始,任务成功率显著提升,从低于5%上升到超过15%,但与人工辅助重建相比仍有较大差距。我们进一步发现,看起来更好的世界可能效果更差:更好的视觉保真度并不总是带来更高的任务成功率。这呼应了生成模型中感知真实性与事实正确性之间更广泛的差距。
英文摘要
Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5\% to over 15\%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.
发表机构
- Aether AI
- University of California, San Diego(加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。