RoGSW4RLD:用于机器人世界模型推演的前馈4D高斯提升
RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts
浏览论文内容
中文总结 AI 辅助
RoGSW4RLD通过两阶段前馈架构将多相机视频推演提升为统一4D高斯场,显著提升新视角PSNR、降低深度误差和位移误差。
中文摘要 AI 辅助
动作条件视频世界模型从多个相机预测未来的机器人交互,但其输出仍然是分散的视频集合,而非可在不同视角和时间上查询的共享度量场景。虽然现有的4D重建方法为空间化这些预测提供了途径,但独立重建并合并每个相机流无法强制跨视角一致性。这一限制在结合移动机器人搭载相机与固定外部视角时尤为不利。为解决此问题,我们提出RoGSW4RLD,一种前馈框架,将同步的多相机推演提升为统一的、可时间查询的度量4D高斯场。RoGSW4RLD不学习单独的几何转换模型,而是直接重建现有世界模型生成的视觉未来。其核心创新在于两阶段架构:第一阶段通过融合跨视角证据与机器人特定的关节几何和运动学,联合形成度量4D场;第二阶段在严格保持初始时间位移的同时,细化场的几何和外观。在256个保留的DROID片段上评估,RoGSW4RLD显著优于带校准合并的逐相机重建,新视角PSNR提高2.15 dB,深度AbsRel降低47%,机器人位移误差降低61%。这些稳健的提升扩展到动作条件的Cosmos 3推演,表明预测的视频未来可成功转换为一致、可空间查询的4D度量表示。
英文摘要
Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently reconstructing and merging each camera stream fails to enforce cross-view consistency. This limitation is particularly detrimental when combining moving robot-mounted cameras with fixed external views. To address this, we introduce RoGSW4RLD, a feed-forward framework that lifts synchronized multi-camera rollouts into a unified, time-queryable metric 4D Gaussian field. Rather than learning a separate geometric transition model, RoGSW4RLD directly reconstructs the visual future generated by existing world models. Its core innovation is a two-stage architecture: Stage 1 jointly forms the metric 4D field by fusing cross-view evidence with robot-specific articulated geometry and kinematics, while Stage 2 refines the field's geometry and appearance while strictly preserving the initial temporal displacements. Evaluated on 256 held-out DROID episodes, RoGSW4RLD significantly outperforms camera-wise reconstruction with calibrated merging, improving novel-view PSNR by 2.15 dB, reducing depth AbsRel by 47%, and lowering robot displacement error by 61%. These robust gains extend to action-conditioned Cosmos 3 rollouts, demonstrating that predicted video futures can be successfully translated into consistent, spatially queryable 4D metric representations.
发表机构
- Korea University(高丽大学)
- Dongguk University(东国大学)
机构由 AI 辅助整理,请以论文原文为准。