VPTwin:用于机器人操作规划的真实-仿真-真实视频预测
VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning
浏览论文内容
中文总结 AI 辅助
VPTwin提出真实-仿真-真实视频预测框架,利用仿真孪生和VLM重建,减少物理幻觉并提升机器人操作规划性能。
中文摘要 AI 辅助
尽管基于动作条件的视频预测为机器人提供了一个直观的世界模型,但纯数据驱动的预测器在长时程推演中常常遭受累积误差和物理上不合理的幻觉,严重削弱了下游的动作规划。我们提出了VPTwin,一个真实-仿真-真实的视频预测框架,该框架利用真实同步的仿真孪生来锚定真实世界的未来预测。对于目标操作任务,视觉语言模型(VLM)从真实演示片段中重建一个可执行的数字孪生。为了适应未观测物理属性的不适定估计,Isaac Sim在候选动作轨迹下,跨随机化物理配置模拟多个前向动力学推演。利用这些推演作为上下文参考,VPTwin协调两个域,利用仿真动力学强制物理合理性,同时从真实视频中捕捉未建模的接触交互。此外,我们建立了一个基于VPTwin的预测规划循环,以视觉验证VLM提出的动作,并引导可靠的真实世界执行。评估表明,在视频预测期间物理幻觉显著减少,操作规划性能显著提升。
英文摘要
While action-conditioned video prediction provides an intuitive world model for robotics, purely data-driven predictors often suffer from compounding errors and physically implausible hallucinations in long-horizon rollouts, severely undermining downstream action planning. We propose VPTwin, a Real-Sim-Real video prediction framework that anchors real-world future prediction using real-synchronized simulation twins. For a target manipulation task, a VLM reconstructs an executable digital twin from a real demonstration episode. To accommodate the ill-posed estimation of unobserved physical properties, Isaac Sim simulates multiple forward dynamic rollouts across randomized physical configurations under candidate action trajectories. Using these rollouts as in-context references, VPTwin harmonizes both domains, using simulation dynamics to enforce physical plausibility while capturing unmodeled contact interactions from real video. Furthermore, we establish a predictive planning loop using VPTwin to visually verify VLM-proposed actions and guide reliable real-world execution. Evaluations show substantial reductions in physical hallucinations during video prediction and marked improvements in manipulation planning performance.