WorldLine:用于机器人操作的动作驱动视觉仿真
WorldLine: Action-Driven Visual Simulation for Robotic Manipulation
- The Hong Kong University of Science and Technology(香港科技大学)
- Joy Future Academy(京东探索研究院)
- Peking University(北京大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
WorldLine是一种动作驱动的视觉模拟器,利用超万小时无动作视频和两千小时多形态动作轨迹解耦动态学习与动作接地,实现高效策略评估,提升任务成功率。
AI中文摘要:
真实世界中的机器人学习受到收集经验和评估候选行为成本的限制。视频生成模型为视觉模拟器提供了可扩展的基础,使其能够在物理执行之前预测动作结果。然而,这些模型往往更倾向于视觉合理性,而非准确的动作遵循和连贯的机器人-物体动态,而动作条件模拟器则依赖于稀缺的、特定于具体形态的数据,这些数据难以在不兼容的控制空间之间共享。我们提出了WorldLine,一种动作驱动的视觉模拟器,它将可迁移的动态学习与异构动作接地解耦。WorldLine从超过10,000小时的无动作机器人视频中学习操作动态,并利用来自十多种形态的超过2,000小时的动作轨迹对其进行接地。一种图像空间动作表示提供了跨形态的共享控制接口,而多视图和带关系正则化的失败增强训练则改善了交互敏感预测。以机器人为中心的少步蒸馏实现了高效的因果 rollout,同时保留了动作关键运动。在留出和域外设置中,WorldLine保持了强大的视觉质量和机器人运动一致性;在失败轨迹上,它相对于最强基线将机器人掩码IoU提高了0.1626。它在RoboTwin和AgiBot上以74%的平均准确率预测轨迹成功,比最强基线高出一个百分点。在没有RoboTwin训练或适应的情况下,其rollout相对于直接策略执行将任务成功率提高了最多21.4个百分点。这些能力共同使WorldLine成为用于策略评估和具身规划的可扩展且高效的视觉模拟器。更多结果可在项目页面获取。
英文摘要:
Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \href{https://zhengsh123.github.io/WorldLine/}{project page}.