arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05834cs.AI

学习部分可观测条件下具身推理的反事实世界模型

Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

Todd Y. Zhou, Daniel Zhang

AI总结:

针对部分可观测下世界模型的反事实坍缩问题,提出CLWM模型,结合循环编码器与对比目标提升规划成功率,并引入反事实可分性指标以审计编码器可操作性。

AI中文摘要:

世界模型承诺了一条通往具身智能的通用路径:先学习预测性动态,然后据此进行推理、规划和行动。日益增多的是,此类模型底层的表示是在大规模视频、交互和多模态语料库上预训练的,这引发了一个仅靠预测质量无法回答的问题:学习到的表示何时真正可操作?我们识别出一种称为反事实坍缩的失败模式:模型预测出视觉上合理的未来,却无法区分具有不同行为后果的干预。只要表示是针对感知相似性而非干预结构进行优化的,就会出现这种情况,而这正是大多数大规模预训练编码器的学习目标。我们引入了反事实潜在世界模型(CLWM),它结合了循环信念状态编码器、动作条件潜在动态以及对比反事实目标,即使干预产生的观测看起来相似,也能区分由不同干预导致的未来。在遮挡操作、别名导航和长视界操作中,CLWM 相比最强基线提高了规划成功率(遮挡推动从 65.1% 提升至 74.6%,别名迷宫从 67.3% 提升至 78.9%),并减少了剥削性规划失败(延迟厨房从 18.4% 降至 9.7%),消融实验将增益归因于困难反事实负样本,尤其是感知别名负样本。最后,我们的反事实可分性指标在五种基线模型类别中与规划成功率高度相关(r ≥ 0.94),且该指标与表示无关:给定干预-结果标签,它可以在规划器信任任何编码器(预训练或从头训练)之前对其进行审计。我们尚未在大规模预训练编码器上测量该指标。在此,我们为从头训练的世界模型建立了该指标及其与规划成功率的关系。

英文摘要:

World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the representations beneath such models are pretrained on large-scale video, interaction, and multimodal corpora, which raises a question prediction quality alone cannot answer: when is a learned representation actually actionable? We identify a failure mode we call counterfactual collapse: a model predicts visually plausible futures while failing to distinguish interventions with different behavioral consequences. This arises whenever a representation is optimized for perceptual similarity rather than intervention structure, which is precisely the objective under which most large-scale pretrained encoders are learned. We introduce Counterfactual Latent World Models (CLWM), which combine a recurrent belief-state encoder, action-conditioned latent dynamics, and a contrastive counterfactual objective that separates futures induced by distinct interventions even when their observations look alike. Across occluded manipulation, aliased navigation, and long-horizon manipulation, CLWM improves planning success over the strongest baseline (65.1% $\to$ 74.6% on Occluded Push and 67.3% $\to$ 78.9% on Aliased Maze) and reduces exploitative planning failures (18.4% $\to$ 9.7% on Deferred Kitchen), with ablations attributing the gains to hard counterfactual negatives, especially perceptual-alias negatives. Finally, our counterfactual separability metric, which tracks planning success across the five baseline model classes ($r \ge 0.94$), is representation-agnostic: given intervention-outcome labels, it can audit any encoder, pretrained or trained from scratch, before a planner trusts it. We do not yet measure it on large-scale pretrained encoders. Here we establish the metric and its relationship to planning success for world models trained from scratch.

↑