发表机构
Seoul National University; Ewha Womans University(首尔大学; 梨花女子大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对JEPA世界模型中各向同性正则化导致规划成本与任务不一致的问题,提出AnisoWM,采用可学习对角协方差正则化,在四个视觉控制环境中均提升规划成功率。
AI 中文摘要
潜在世界模型在表示空间中学习动作条件下的动力学,并通常通过到目标表示的欧几里得距离来对候选动作进行评分。联合训练通常会正则化表示以防止坍缩,但由此产生的表示几何也决定了规划过程中终端误差的加权方式。我们表明,准确的预测和非坍缩的表示并不能保证任务对齐的潜在规划成本:各向同性高斯正则化可能诱导一种几何结构,使得对可行结果的排序与任务成本不同。为解决这一不匹配问题,我们引入了带有$\Lambda$Reg的AnisoWM,它将固定的各向同性高斯目标替换为在固定迹和各向异性约束下的可学习对角协方差。预测目标、预测器架构和欧几里得规划器保持不变;该目标仅在训练期间使用。我们的分析刻画了预测驱动的目标方差分配、其对训练分布的依赖性,以及诱导度量减少规划遗憾的条件。在四个视觉控制环境中,AnisoWM在全部四个环境中相比LeWorldModel提高了规划成功率。其潜在规划成本与任务结果的一致性也更好。项目网站:此https URL
英文摘要
Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with $Λ$Reg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/