发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在潜在世界模型中分离世界效应与动作的问题,提出DWM框架,通过辅助世界头和正交性约束实现预测转换的显式加法分解,在构建的W变体基准测试中取得更好效果。
AI 中文摘要
潜在世界模型是现代基于模型控制的基础,但当前基于动作条件的公式用单一、无差别的目标监督下一个潜在转换,迫使整体学习信号吸收状态变化的每个来源。在现实世界中,转换源于两个异质源:由智能体引起的动作驱动组件和动作不变的世界效应。我们引入DWM(分解世界模型),通过辅助世界头增强潜在世界模型的预测器,由归一化世界对比目标正则化以实现动作不变,同时原始预测头通过正交性约束与之耦合。为评估DWM在持续世界效应下的性能,构建了三个标准控制基准的W变体。DWM在平坦对应物上与强大基线匹配,在W变体的CEM规划成功率上平均绝对提高了13.1%。
英文摘要
Latent world models underpin much of modern model-based control, yet current action-conditioned formulations supervise the next-latent transition with a single, undifferentiated target, forcing a monolithic learning signal to absorb every source of state change. In real world, however, transitions arise from two heterogeneous sources: an action-driven component induced by the agent, and an action-invariant world effect -- the change that would still occur under a null action, dictated by the environment's intrinsic dynamics (e.g., gravity-driven sliding, inertia, contact rebound, and persistent drift). Fusing them into a single target entangles the two inside the latent transition, prevents the model from attributing observed changes to their underlying causes, and undermines the transferability of the learned dynamics. We introduce DWM (Decomposed World Model), a supervision-level framework that operationalizes this decomposition. DWM augments the predictor of a latent world model with an auxiliary world head, regularized by a normalized world-contrastive objective to be action-invariant, while the original pred head is coupled to it via an orthogonality constraint; together, the two signals induce an explicit additive decomposition of the predicted transition into an action-invariant and a complementary action-driven component, without altering the underlying architecture or inference pipeline. To evaluate DWM under persistent world effects, we construct W-variants of three standard control benchmarks -- PushT-W, Reacher-W, and TwoRoom-W -- each instantiating a distinct action-invariant dynamic. DWM matches strong baselines on the flat counterparts and delivers a mean absolute improvement of 13.1% in CEM planning success across the W-variants.