发表机构
Nanyang Technological University; University of Oxford; Princeton University; Google DeepMind(南洋理工大学; 牛津大学; 普林斯顿大学; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出可控程序式4D环境生成系统4DSynth,可将自然语言等输入转化为可编辑4D环境,并构建交互式导航基准4DSynth-Nav,为具身智能体开发评估提供实用基础。
AI 中文摘要
具身智能体需要视觉多样、物理可交互且随时间变化的环境。程序式模拟器可生成大量交互场景集合,近期的4D生成器能产出极具吸引力的视觉动态效果。然而,将这些特性整合到单一环境中仍需大量人工工作,且生成结果极少具备足够的可编辑性或可控性以实现大规模复用。我们提出4DSynth,这是一种可控程序式系统,可将自然语言描述、蓝图掩码或单张照片转化为具备显式几何结构、动画角色、无碰撞轨迹及物理就绪模拟状态的可编辑4D环境。多条场景路径共享同一几何基础表示,因此同一流程可处理动画、相机规划、渲染及任务生成。为验证完整流程,我们构建了4DSynth-Nav,这是一个完全由4DSynth程序式场景生成的交互式导航基准。在三个难度层级上评估的两个视觉语言模型均未能完成多数任务,并在早期子任务后弃权(不执行)。生成这些环境的相同程序式可控性使每个失败案例可复现,且每个难度轴可独立调优。本文既提出了可控生成流程,也提供了其支持的可扩展基准,为开发和评估具身智能体提供了实用基础。
英文摘要
Embodied agents need environments that are visually diverse, physically interactive, and changing over time. Procedural simulators can generate large interactive scene collections, and recent 4D generators produce compelling visual dynamics. Combining these properties in one environment, however, still demands extensive manual effort, and the result is rarely editable or controllable enough to reuse at scale. We present 4DSynth, a controllable procedural system that turns a natural-language description, a blueprint mask, or a single photograph into an editable 4D environment with explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation state. Multiple scene routes share one geometry-grounded representation, so the same pipeline handles animation, camera planning, rendering, and task generation. To validate the full pipeline, we construct 4DSynth-Nav, an interactive navigation benchmark generated entirely from 4DSynth's procedural scenes. Two vision-language models evaluated across three difficulty tiers both fail the majority of tasks and stall after early subtasks. The same procedural controllability that produces these environments also makes each failure reproducible and each difficulty axis independently tunable. This paper presents both a controllable generation pipeline and the scalable benchmark it enables, offering a practical foundation for developing and evaluating embodied agents.