AI 中文总结
该研究针对动作可控世界模型的统计偏差问题,提出CoCo反事实一致性框架,结合ARC、DE评估指标及Mini-SSMB数据集,在多项任务上提升了动作可控性与视频预测性能。
AI 中文摘要
动作可控世界模型旨在预测智能体的作用下视觉环境如何演变。然而,未来帧往往仅通过视觉惯性和重复的运动模式就具有高度可预测性,这形成了一种捷径:模型可以通过利用统计偏差来拟合数据,而无需使其可见动态有意义地依赖于动作。因此,不同动作可能产生相似的未来,即使在零动作下运动也可能持续存在。关键问题是如何减少对统计捷径的依赖,使其不会主导动作条件下的预测。我们认为,动作控制不仅仅需要注入动作特征,还需要在动作和观测的反事实变化下强制执行一致性。基于这一见解,我们引入了CoCo(Counterfactual Consistency,反事实一致性)框架,通过两个互补约束来增强动作可控性:多步反事实一致性约束参考、逆动作和零动作的滚动,而动作空间反事实一致性则在镜像场景和变换动作下强制执行一致的预测。这些约束共同减少了对统计捷径的依赖,避免其替代依赖于动作的动态。我们还引入了动作响应一致性(ARC)和漂移能量(DE)来评估动作可控性,同时引入了Mini-SSMB用于同状态多动作反事实评估。在Mini-SSMB上,我们的完整模型实现了ARC_inv为0.412、ARC_ref为0.483,且相对于基线将DE降低了17.07%;在VP2视觉规划任务中,它达到了SOTA模型中最高的平均成功率,为73.1%。在BAIR和RoboNet上的进一步实验表明,这些提升在保留视频预测质量的同时,可跨模型设置迁移。
英文摘要
Action-conditioned world models aim to predict how visual environments evolve under an agent's actions. Yet future frames are often highly predictable from visual inertia and recurring motion patterns alone. This creates a shortcut: models can fit the data by exploiting statistical biases without making their visible dynamics meaningfully depend on the action. As a result, different actions may produce similar futures, while motion may persist even under zero action. The key question is how to reduce reliance on statistical shortcuts from dominating action-conditioned prediction. We argue that action control requires more than injecting action features; it requires enforcing consistency under counterfactual changes to actions and observations. Based on this insight, we introduce CoCo, a Counterfactual Consistency framework to enhance action controllability through two complementary constraints. Multi-step counterfactual consistency constrains reference, inverse-action, and zero-action rollouts, while action-spatial counterfactual consistency enforces consistent predictions under mirrored scenes and transformed actions. Together, they reduce reliance on statistical shortcuts from substituting for action-dependent dynamics. We further introduce Action Response Consistency (ARC) and Drift Energy (DE) to assess action controllability, together with Mini-SSMB for same-state, multi-action counterfactual evaluation. On Mini-SSMB, our full model achieved ARC_inv of 0.412 and ARC_ref of 0.483, while reducing DE by 17.07% relative to the baseline. On VP2 visual planning, it achieves the highest average success rate among SOTA models, at 73.1%. Experiments on BAIR and RoboNet further show that these gains preserve video prediction quality and transfer across model settings.