发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出在世界模型潜在空间中对潜在扰动建模,通过保形预测校准不确定性集并采用博弈论优化,实现鲁棒决策,在仿真和Franka机械臂实验中显著降低失败率。
AI 中文摘要
在本文中,我们研究了世界模型(WMs)潜在空间中的鲁棒决策。鲁棒优化是一个数学框架,在给定明确指定的动力学和物理上有意义的扰动的情况下,机器人可以选择即使在最坏情况扰动下仍然有效的动作。然而,将此原则应用于WMs的学习潜在空间引入了一个根本性挑战:因为WMs具有完全学习的、从高维观测中推断出的状态空间和动力学,如何定义能忠实表示底层系统中不确定性的潜在空间扰动尚不清楚。我们的关键思想是将潜在空间扰动建模为对学习到的潜在动力学的扰动,该扰动诱导悲观但合理的转移。具体来说,我们通过结合动力学感知的相似性度量(捕捉合理转移)和分布外检测(排除不合理的潜在状态)来构建一组合理的潜在动力学。我们使用保形预测校准这个关于潜在动力学的不确定性集,确保由潜在扰动诱导的WM想象保持合理而不过度悲观。然后,我们通过博弈论优化联合优化鲁棒机器人动作和最坏情况潜在扰动。我们利用这种潜在空间鲁棒优化来强化策略引导,考虑两种范式:潜在安全过滤和生成控制策略的采样-验证引导。我们的受控仿真实验表明,我们的潜在扰动能够在WM潜在空间中直接实现鲁棒决策,而使用Franka机械臂的硬件实验表明,对潜在扰动进行建模能够实现鲁棒策略引导,在安全过滤中减少70%的失败,在基于采样的策略引导中减少54%的失败。项目网站:此https URL。
英文摘要
In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, applying this principle to the learned latent space of WMs introduces a fundamental challenge: because WMs have fully learned state spaces and dynamics inferred from high-dimensional observations, it is unclear how to define latent-space disturbances that faithfully represent uncertainty in the underlying system. Our key idea is to model a latent-space disturbance as a perturbation to the learned latent dynamics that induces pessimistic but plausible transitions. Specifically, we construct a set of plausible latent dynamics by combining a dynamics-aware similarity metric that captures plausible transitions with out-of-distribution detection that excludes implausible latent states. We calibrate this uncertainty set over latent dynamics using conformal prediction, ensuring that WM imaginations induced by the latent disturbance remain plausible without becoming overly pessimistic. We then jointly optimize robust robot actions and the worst-case latent disturbances through game-theoretic optimization. We leverage this latent-space robust optimization to robustify policy steering, considering two paradigms: latent safety filtering and sample-and-verify steering of a generative control policy. Our controlled simulation experiments show that our latent disturbance enables robust decision-making directly in WM latent spaces, and hardware experiments with a Franka manipulator show that modeling latent disturbances enables robust policy steering, reducing failures by 70% in safety filtering and 54% in sampling-based policy steering. Project website: https://junwon.me/LatentDisturbance/.