当世界模型说谎:错误想象下的自适应安全分析
When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations
浏览论文内容
中文总结 AI 辅助
针对世界模型预测错误导致的安全估计过度自信问题,提出一种自适应潜在安全过滤器,利用观测到的模型误差构建不确定性集合并悲观评估安全性,在仿真和硬件实验中减少了失败并保持任务完成。
中文摘要 AI 辅助
世界模型为高维机器人系统中的安全推理提供了强大的基础,但它们也可能出错:其预测可能存在偏差、校准不良或自信地错误。这给潜在空间安全过滤器带来了核心挑战,这些过滤器通常在世界模型的动力学上学习Hamilton-Jacobi安全价值函数。如果世界模型不正确,所得的价值函数可能继承其错误并产生过度自信的安全估计。现有的潜在安全过滤器通常依赖集成分歧或价值目标一致性残差等辅助信号进行自适应,但这些信号即使在世界模型的预测偏离观测时也可能保持较小。我们提出了一种自适应潜在安全过滤器,利用直接观测到的世界模型误差来校准安全推理。我们的方法使用自适应共形推断,根据预测潜在状态与观测推断潜在状态之间的差异构建在线不确定性集合,然后通过在这些集合上最小化学习到的价值函数来悲观地评估安全性。这使得过滤器在世界模型准确时保持最小保守性,而在观测揭示模型失配时变得更加谨慎。我们为自适应不确定性半径提供了有限时间覆盖保证。通过仿真和硬件实验,我们表明相对于最先进的潜在安全过滤器,我们的方法显著减少了失败,同时保持了任务完成率。
英文摘要
World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions on the dynamics of a world model. If the world model is incorrect, the resulting value function can inherit its errors and produce overconfident safety estimates. Existing latent safety filters often rely on auxiliary signals such as ensemble disagreement or value-target consistency residuals for adaptation, but these signals can remain small even when the world model's predictions deviate from observations. We propose an adaptive latent safety filter that calibrates safety reasoning using directly observed world-model error. Our method uses Adaptive Conformal Inference to construct online uncertainty sets from discrepancies between predicted and observation-inferred latent states, then evaluates safety pessimistically by minimizing the learned value function over these sets. This allows the filter to remain minimally conservative when the world model is accurate, while becoming more cautious when observations reveal model mismatch. We provide a finite-time coverage guarantee for the adaptive uncertainty radius. Through simulation and hardware experiments, we show that our method significantly reduces failures relative to state-of-the-art latent safety filters while preserving task completion.
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。