NeoWorld-Pro:从单目图像编程交互式场景以实现具身仿真
NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation
- Shanghai Jiao Tong University(上海交通大学)
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对具身AI中图像转仿真场景的物理与交互性不足问题,提出NeoWorld-Pro框架,通过MLLMs将单目图像转为可执行程序,结合物理在环机制优化,性能优于现有方法并支持复杂下游任务。
AI中文摘要:
具身AI的发展需要能忠实反映现实世界的高质量仿真资产。然而,当前的图像转URDF方法因缺乏物理 grounding(接地)和场景级交互性,将原始视觉观测转换为可用于仿真的场景仍具挑战性。我们提出NeoWorld-Pro,这一框架将单目场景重建重新表述为交互式3D环境的程序化编程。利用多模态大语言模型(MLLMs)的零样本推理和代码合成能力,NeoWorld-Pro将单张RGB图像转换为可执行程序,这些程序指定了物体的几何结构、关节运动特性和物理属性。随后,物理在环机制通过在物理引擎中验证生成程序的执行来迭代优化这些程序,确保关节运动符合物理规律、物体组成与交互合理,以及空间关系准确。实验表明,NeoWorld-Pro的性能优于开环方法和现有的单目重建方法,同时支持稳定堆叠、精细操作等复杂下游任务。
英文摘要:
The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the real world. However, transforming raw visual observations into simulation-ready scenes remains challenging due to the lack of physical grounding and scene-level interactivity in current image-to-URDF methods. We propose NeoWorld-Pro, a framework that reformulates monocular scene reconstruction as procedural programming for interactive 3D environments. Leveraging the zero-shot reasoning and code synthesis capabilities of MLLMs, NeoWorld-Pro converts a single RGB image into executable programs specifying object geometry, articulation, and physical properties. A physics-in-the-loop mechanism then iteratively refines the generated programs by validating their execution in a physics engine, enforcing physically plausible articulations, valid object compositions and interactions, and accurate spatial relationships. Experiments show that NeoWorld-Pro outperforms open-loop and prior monocular reconstruction methods, while enabling complex downstream tasks such as stable stacking and fine-grained manipulation.