DepthWorld:用于机器人操作的3D世界模型
DepthWorld: 3D World Model for Robot Manipulation
浏览论文内容
中文总结 AI 辅助
本文提出DepthWorld,一种基于稳定视频扩散的3D世界模型,通过校准的DROID-3D数据集和联合深度-RGB预测,实现机器人操作的精确几何建模与策略评估。
中文摘要 AI 辅助
世界模型为机器人学提供了传统模拟器的数据驱动替代方案,其应用涵盖策略评估、改进和规划。所有这些用途都依赖于忠实的3D几何结构,然而当前基于视频的世界模型仅使用RGB数据进行训练,生成的展开结果在逐帧层面看似正确,但无法组合成一致的3D世界。弥合这一差距需要在两个方面取得进展:用于操作的大规模3D监督,以及能够在不过度干扰预训练视频先验的情况下吸收这种监督的架构。我们引入了一个校准流程,该流程将学习到的立体深度与联合因子图相结合,汇集了从同一物理机器人收集的所有片段,以恢复其共享的运动学参数以及每个场景的外部参数。将该流程应用于DROID数据集,我们生成了DROID-3D,这是一个经过校准的3D数据集,提供密集的度量深度和重新校准的多视角外部参数(在90%的外部相机片段中实现了<0.7像素的重投影误差)。随后,我们训练了DepthWorld,这是一个基于稳定视频扩散的世界模型,通过空间潜在平铺联合预测多视角RGB和深度,同时保持预训练变分自编码器(VAE)不变。在相同的训练预算下,深度监督使RGB预测本身比仅使用RGB的基线提高了+1.48 dB的PSNR,同时为下游几何推理提供了准确的度量深度。
英文摘要
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving <0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.
发表机构
- Czech Technical University in Prague(捷克理工大学)
机构由 AI 辅助整理,请以论文原文为准。