机器人世界模型真的会遵循动作吗?面向策略学习的动作条件生成诊断与对齐
Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning
浏览论文内容
中文总结 AI 辅助
本文针对动作条件世界模型的动作遵循问题,提出WorldEcho诊断工具与WorldSync改进方法,经实验验证WorldSync可提升动作遵循性能,作为可靠模拟器助力策略改进并提高任务成功率。
中文摘要 AI 辅助
动作条件世界模型正日益被用作策略评估与改进的学习型模拟器,但其有效性依赖一个未经验证的假设:生成的未来结果能忠实地反映任意有效动作。现有基准通常局限于专家演示,对专家外动作的遵循情况评估不足。为解决这一缺口,我们引入WorldEcho,该工具通过视觉完整性和SE(3)轨迹对齐,在更广泛的动作分布上探测动作遵循情况。我们的诊断显示,当前世界模型能合理执行专家动作,但在多样化的专家外轨迹上表现不佳,要么忽略指令动作,要么生成视觉上无效的 rollout。我们进一步提出WorldSync,它从三个互补维度强化动作遵循:分布覆盖、表征 grounding、干预效应对齐。它扩大了动作后果的训练分布,通过动作强制专家(Action-Forcing Expert)将中间视频表征与动作诱导的机器人动力学建立 grounding,还使动作干预下的预测变化与真实未来的对应变化对齐。在RoboTwin基准和真实机器人任务上的实验表明,WorldSync 提升了WorldEcho指标,可作为更可靠的模拟器用于迭代策略改进,使策略达到更高的成功率。
英文摘要
Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-expert action following inadequately evaluated. To address this gap, we introduce WorldEcho, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. Our diagnosis shows that current world models reasonably execute expert actions but struggle with diverse off-expert trajectories, either ignoring the commanded actions or producing visually invalid rollouts. We further propose WorldSync, which strengthens action following along three complementary axes: distributional coverage, representational grounding, and intervention-effect alignment. It broadens the training distribution over action consequences, grounds intermediate video representations in action-induced robot dynamics through an Action-Forcing Expert, and aligns predicted changes under action interventions with the corresponding changes in ground-truth futures. Experiments on RoboTwin benchmarks and real-robot tasks show that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to achieve higher success rates.
发表机构
- Peking University(北京大学)
- Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
- New York University(纽约大学)
- University of Electronic Science and Technology of China(电子科技大学)
- Nanyang Technological University(南洋理工大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。