arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24885cs.ROcs.CV

机器人世界模型真的会遵循动作吗?面向策略学习的动作条件生成诊断与对齐

Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

Sixiang Chen, Jiaming Liu, Jixian Wu, Yichen Guo, Tinghao Wang, Siyuan Qian, Hao Chen, Jiajun Cao, Jian Tang, Shanghang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对动作条件世界模型的动作遵循问题,提出WorldEcho诊断工具与WorldSync改进方法,经实验验证WorldSync可提升动作遵循性能,作为可靠模拟器助力策略改进并提高任务成功率。

中文摘要 AI 辅助

动作条件世界模型正日益被用作策略评估与改进的学习型模拟器,但其有效性依赖一个未经验证的假设:生成的未来结果能忠实地反映任意有效动作。现有基准通常局限于专家演示,对专家外动作的遵循情况评估不足。为解决这一缺口,我们引入WorldEcho,该工具通过视觉完整性和SE(3)轨迹对齐,在更广泛的动作分布上探测动作遵循情况。我们的诊断显示,当前世界模型能合理执行专家动作,但在多样化的专家外轨迹上表现不佳,要么忽略指令动作,要么生成视觉上无效的 rollout。我们进一步提出WorldSync,它从三个互补维度强化动作遵循:分布覆盖、表征 grounding、干预效应对齐。它扩大了动作后果的训练分布,通过动作强制专家(Action-Forcing Expert)将中间视频表征与动作诱导的机器人动力学建立 grounding,还使动作干预下的预测变化与真实未来的对应变化对齐。在RoboTwin基准和真实机器人任务上的实验表明,WorldSync 提升了WorldEcho指标,可作为更可靠的模拟器用于迭代策略改进,使策略达到更高的成功率。

英文摘要

Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-expert action following inadequately evaluated. To address this gap, we introduce WorldEcho, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. Our diagnosis shows that current world models reasonably execute expert actions but struggle with diverse off-expert trajectories, either ignoring the commanded actions or producing visually invalid rollouts. We further propose WorldSync, which strengthens action following along three complementary axes: distributional coverage, representational grounding, and intervention-effect alignment. It broadens the training distribution over action consequences, grounds intermediate video representations in action-induced robot dynamics through an Action-Forcing Expert, and aligns predicted changes under action interventions with the corresponding changes in ground-truth futures. Experiments on RoboTwin benchmarks and real-robot tasks show that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to achieve higher success rates.

发表机构

  • Peking University(北京大学)
  • Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
  • New York University(纽约大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Nanyang Technological University(南洋理工大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑