NavGen:视觉生成模型作为具身3D导航的可扩展数据引擎
NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation
浏览论文内容
中文总结 AI 辅助
针对具身3D导航数据在模拟与真实之间的权衡,提出利用高保真视觉生成模型作为数据引擎,构建NavGen文本到视频流水线,生成约40万导航片段,训练模型在真实飞行实验中达到75%成功率。
中文摘要 AI 辅助
通用机器人模型越来越依赖于大规模和多样化的数据集。然而,对于具身3D导航而言,现有数据源面临一个根本性的权衡:模拟数据可以大规模生成,但常常存在视觉上的模拟到现实差距,而真实世界的飞行数据提供了逼真的观测,但收集成本高昂。本文研究了另一个方向:利用高保真视觉生成模型作为具身3D导航的可扩展数据引擎。我们引入了NavGen,一个文本到视频的数据生成流水线,可在室内和室外场景中生成多样化的视觉语言导航(VLN)片段。我们还提出了一种风格多样化方法,用于扩展难以且成本高昂收集的长尾数据。由此产生的数据集包含约40万个导航片段。我们在多个指标上对现有无人机导航数据集评估了我们的数据集,发现基于我们数据训练的模型通常随规模扩大而改善,优于基于现有数据集训练的模型。为了验证真实世界的可迁移性,我们将训练后的模型部署在世界行动模型范式中进行真实飞行实验。最终模型在不同导航任务和环境中的成功率达到75%。
英文摘要
General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, whereas real-world flight data provide realistic observations but are costly to collect. This paper studies another direction: the use of high-fidelity visual generative models as scalable data engines for embodied 3D navigation. We introduce NavGen, a text-to-video data generation pipeline that produces diverse vision-language navigation (VLN) episodes across indoor and outdoor scenes. We also propose a style-diversification method that scales up long-tail data that are difficult and costly to collect. The resulting dataset contains approximately 400K navigation episodes. We evaluate our dataset against existing UAV navigation datasets across multiple metrics, and find that the model trained on our data generally improves with scale, outperforming those trained on existing datasets. To validate real-world transferability, we deploy the trained model in world-action-model paradigm to real-world flying experiments. The final model achieves a 75\% success rate across different navigation tasks and environments.
发表机构
- Zhejiang University(浙江大学)
- Differential Robotics
机构由 AI 辅助整理,请以论文原文为准。