发表机构
MIT; University of Toronto(麻省理工学院; 多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EnvDreamer利用大型语言和视觉语言模型自动生成Unreal Engine 5虚拟环境,用于具身AI训练,无需人工监督即可在多个基准上取得有竞争力结果,并发布20k环境数据集。
AI 中文摘要
大型数据集和高容量模型加速了视觉和语言领域的进展。这项工作引入了一个平台,旨在为具身学习、世界模型和机器人技术带来类似的收益。我们提出了EnvDreamer,一个利用大型语言模型和视觉语言模型为具身AI和机器人训练生成Unreal Engine 5环境的框架。EnvDreamer能够采样大规模、多样化、可交互、可定制且通过验证器的虚拟环境,用于导航、交互和操作任务的训练与评估。我们用大量生成的场景和简单基线展示了该平台。在EnvDreamer生成的环境上训练的策略,无需显式映射或人工任务监督,即在多个涵盖导航、重排和操作的具身基准上取得了有竞争力的结果。EnvDreamer还支持图像条件重建,用于真实到模拟的研究。最后,我们发布了EnvDreamer-20k,一个包含20,000个通过验证器的环境数据集,附带任务程序、场景图、轨迹和元数据,以支持可复现的基准测试。
英文摘要
Large datasets and high capacity models have accelerated progress in vision and language. This work introduces a platform aimed at bringing comparable gains to embodied learning, world models, and robotics. We present EnvDreamer, a framework that uses large language and vision language models to generate Unreal Engine 5 environments for embodied AI and robot training. EnvDreamer enables sampling of large, diverse, interactive, customizable, and validator passed virtual environments for training and evaluation across navigation, interaction, and manipulation. We illustrate the platform with a large set of generated scenes and simple baselines. Policies trained on EnvDreamer generated environments, without explicit mapping or human task supervision, achieve competitive results on multiple embodied benchmarks spanning navigation, rearrangement, and manipulation. EnvDreamer also supports image-conditioned reconstruction for real-to-sim studies. Finally, we release EnvDreamer-20k, a dataset of 20,000 validator passed environments with task programs, scene graphs, trajectories, and metadata to support reproducible benchmarking.