构建世界模型的预训练数据:基于Unreal Engine的动作条件视频生成流水线
Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation
浏览论文内容
中文总结 AI 辅助
该研究开发了基于Unreal Engine的分布式流水线,用于生成动作条件多视角合成视频,为世界模型提供大规模预训练数据,已产出大量不同分辨率的视频。
中文摘要 AI 辅助
动作条件视频模型需要大规模视觉数据,且该数据需与随场景变化在时间上对齐的控制信号配对。这类监督信号难以从普通真实世界视频中获取,因为导致每处视觉变化的动作通常是未知的。我们提出一种基于Unreal Engine的大规模合成数据生产流水线,用于生成动作条件多视角视频。为适配实时物理与高质量离线渲染的不同执行需求,该流水线分两个阶段执行轨迹生成与最终渲染:第一阶段在PIE中运行实时物理,将逐帧角色状态、控制输入和相机状态记录至中间轨迹表示;第二阶段在新引擎进程中回放这些轨迹,并使用Movie Render Queue(MRQ)离线渲染。围绕核心功能,我们开发了具备缓存感知任务分区、节点本地槽调度、自动场景筛选、美学与亮度过滤、部分输出恢复、异步上传及持续集群健康监控的分布式生产系统。生产集群包含25台服务器,每台配备8个NVIDIA RTX 5090 GPU。从2384个资源包中,最终保留429个关卡和40个人形角色池用于生产。该流水线已生成2691小时1080p视频与6076小时720p视频。我们描述了系统架构、生产故障中产生的实现决策,以及使用感知质量代理进行世界模型数据筛选的局限性。本报告所述流水线是EchoWM中使用的Unreal Engine合成数据生产组件。
英文摘要
Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is difficult to obtain from ordinary real-world video because the actions that caused each visual change are typically unknown. We present a large-scale synthetic data production pipeline built on Unreal Engine for generating action-conditioned, multi-view video. To accommodate the different execution requirements of real-time physics and high-quality offline rendering, the pipeline executes trajectory generation and final rendering in two stages: Stage I runs real physics in PIE and records per-frame character states, control inputs, and camera states into an intermediate trajectory representation; Stage II replays those trajectories in a new engine process and renders them offline with Movie Render Queue (MRQ). Around this core, we develop a distributed production system with cache-aware task partitioning, node-local slot scheduling, automated scene screening, aesthetic and luminance filtering, partial-output recovery, asynchronous upload, and continuous cluster health monitoring. The production cluster contains 25 servers with eight NVIDIA RTX 5090 GPUs per server. From 2,384 asset packs, 429 levels were retained for production together with a pool of 40 humanoid characters. The pipeline has produced 2,691 hours of 1080p video and 6,076 hours of 720p video. We describe the system architecture, the implementation decisions that emerged from production failures, and the limitations of using perceptual quality proxies for world-model data curation. The pipeline described in this report constitutes the Unreal Engine synthetic-data production component used in EchoWM.
发表机构
- Joy Future Academy, JD(京东探索研究院)
- Tsinghua University(清华大学)
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。