发表机构
Tuojing Intelligence; University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Tsinghua University; University of Science and Technology of China; The Hong Kong University of Science and Technology (Guangzhou); Nanyang Technological University; Beihang University; Carnegie Mellon University; The University of Hong Kong(拓境智能; 中国科学院大学; 中国科学院自动化研究所; 清华大学; 中国科学技术大学; 香港科技大学(广州); 南洋理工大学; 北京航空航天大学; 卡内基梅隆大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GaussianDream++ 是一种高效三维高斯世界建模方法,通过在 VLA 骨干中插入世界状态与预测 token,提升机器人操作性能,在 LIBERO、LIBERO-Plus 及真实机器人实验中均取得显著效果且部署高效。
AI 中文摘要
视觉-语言-动作(VLA)策略在语言条件下的机器人操作方面取得了进展,但动作模仿目标仅为度量三维结构和短时间范围内的物理演化提供了弱监督。几何增强策略主要改进当前场景的 grounding,而预测策略通常在 RGB 或潜在空间中建模未来动力学,可能会产生大量部署成本。GaussianDream 表明,训练时的当前高斯重建和未来高斯预测提供了有效的三维监督,但它基于密集 VGGT/TGE 的前缀共同承载了状态、动力学和动作条件信息。我们提出 GaussianDream++,这是一种紧凑的、策略原生的扩展,直接将 World State Tokens 和 World Prediction Tokens 插入 VLA 骨干网络。仅用于训练的 World Representation Head 将这些 token 解码为共享高斯基元上的当前世界和耦合未来预测,而静态-动态分解保留了持久结构,并将残差运动聚焦于与交互相关的区域。推理时,会移除头部、渲染器、辅助目标和 VGGT/TGE 通路,仅保留 20 个世界 token,无需在线高斯解码或滚动。GaussianDream++ 在 LIBERO 上取得 98.6% 的成绩,在 LIBERO-Plus 上取得 87.8% 的成绩,在相机和布局变化下有明显提升。真实机器人实验相比复现的 π₀.₅ 将平均成功率从 29.2% 进一步提升至 52.5%,同时保持高效的闭环控制。
英文摘要
Vision-Language-Action (VLA) policies have advanced language-conditioned robotic manipulation, yet action-imitation objectives provide only weak supervision for metric 3D structure and short-horizon physical evolution. Geometry-enhanced policies mainly improve current-scene grounding, whereas predictive policies often model future dynamics in RGB or latent spaces and may incur substantial deployment cost. GaussianDream demonstrates that training-time current Gaussian reconstruction and future Gaussian prediction provide effective 3D supervision, but its dense VGGT/TGE-based prefix jointly carries state, dynamics, and action-conditioning information. We present \textbf{\methodname}, a compact, policy-native extension that inserts \textbf{World State Tokens} and \textbf{World Prediction Tokens} directly into the VLA backbone. A training-only \textbf{World Representation Head} decodes these tokens into a Current World and coupled Future Prediction over shared Gaussian primitives, while static--dynamic factorization preserves persistent structure and focuses residual motion on interaction-relevant regions. At inference, the head, renderer, auxiliary objectives, and VGGT/TGE pathway are removed, leaving only 20 world tokens without online Gaussian decoding or rollout. \method achieves \textbf{98.6\%} on LIBERO and \textbf{87.8\%} on LIBERO-Plus, with clear gains under Camera and Layout shifts. Real-robot experiments further improve average success from 29.2\% to 52.5\% over reproduced $π_{0.5}$ while maintaining efficient closed-loop control.
Comments17 pages, 4 figures