arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AD-WM:用于反事实模型预测控制的动作判别世界模型

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao

arXiv 2609.30264首次发表:更新:

发表机构

Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AD-WM提出一种动作判别世界模型,通过残差动力学和动作恢复正则化保留动作信息,显著提升反事实MPC的规划性能,在模拟和真实迁移任务中均大幅提高成功率。

AI 中文摘要

潜在世界模型通常被训练用于预测事实性转移,而模型预测控制(MPC)必须从同一状态比较不同的动作。因此,一个模型可能实现较低的事实预测误差,却难以区分候选动作。我们提出了AD-WM,一种用于反事实MPC的动作判别联合嵌入世界模型。AD-WM将残差潜在动力学与预测器级别的动作恢复正则化相结合,利用逆动力学和基于条件互信息的归一化恢复目标。这两个目标都鼓励规划转移保留动作信息;其辅助头在测试时被丢弃,使MPC保持不变。在OGBench-Cube上,AD-WM将困难起始成功率从3.7%提高到52.0%,超过了匹配的LeWM基线,并在五个模拟环境中的四个环境中提高了平均成功率,优于复现的基线。规划诊断显示,事实预测误差和整库动作排序并不遵循闭环成功排序,而CEM对齐的精英遗憾更接近成功。使用冻结的V-JEPA 2编码器和匹配的DROID后训练,AD-WM还提高了对Franka设置的零样本迁移,将基本拾取和放置成功率从42.2%提高到71.1%,无需实验室特定适应。这些结果表明,用于规划的世界模型应保留反事实选择所需的动作相关差异,而不是仅仅优化事实预测精度。更多视频和代码可在该https URL获取。

英文摘要

Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized recovery objective motivated by conditional mutual information. Both objectives encourage planning transitions to preserve action information; their auxiliary heads are discarded at test time, leaving MPC unchanged. On OGBench-Cube, AD-WM improves hard-start success from 3.7% to 52.0% over a matched LeWM baseline and improves mean success over the reproduced baseline in four of five simulation environments. Planning diagnostics show that factual prediction error and whole-bank action ranking do not follow the closed-loop success ordering, whereas CEM-aligned elite regret tracks success more closely. With a frozen V-JEPA 2 encoder and matched DROID post-training, AD-WM also improves zero-shot transfer to our Franka setup, increasing basic pick-and-place success from 42.2% to 71.1% without lab-specific adaptation. These results suggest that world models for planning should preserve action-dependent differences needed for counterfactual selection, rather than optimize factual prediction accuracy alone. More videos and code are available at https://ad-wm.github.io/.

Comments9 pages, 5 figures, 4 tables. Project page: https://ad-wm.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑