WorldRover:用于带丰富标注的世界探索的可扩展合成视频数据引擎
WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
- Alaya Lab(Alaya实验室)
- The University of Tokyo(东京大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出WorldRover数据引擎,基于Unreal Engine构建可生成带丰富标注的长程探索合成视频的WorldRover-10M数据集,为相关模型提供监督。
中文摘要 AI 辅助
学习生成或重建可探索的世界,需要视频不仅包含RGB信息,还需具备相机运动、场景几何、时间对应关系,以及针对交互模型的控制信号。真实采集可以提供部分此类信号,但密集几何和长程对应关系通常依赖估计或专用仪器。渲染可直接提供这些量,但现有合成资源很少能在同一帧中同时包含这些信息,同时还支持视点和外观的可控变化。我们提出WorldRover,这是一种用于生成艺术家构建环境的带丰富标注的长程探索数据的数据引擎。其核心是WorldRover-Engine,这是一个Unreal Engine管线,可执行并离线渲染分钟级别的路线,同时保留其完整轨迹和场景几何。同一探索过程可通过第一人称、第三人称和360度全景相机,在不同环境状态下重放。利用WorldRover-Engine,我们构建了WorldRover-10M,其序列将RGB与度量深度、相机轨迹以及整个探索过程中从轨迹衍生的动作信号配对。第三人称子集还提供密集光流、带可见性的长程2D/3D点轨迹,以及与相机轨迹不同的角色轨迹。该引擎可从第一人称、第三人称和360度全景视点渲染遍历,在不同环境状态下或使用中性白色材质时,同时保留路线和场景几何。因此,WorldRover将长程世界探索转化为可扩展的数据生成问题,为必须构建、维护和重访可探索世界的连贯表示的模型提供监督。
英文摘要
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.