arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28226cs.CRcs.AI

基于世界模型的具身智能安全性:威胁、防御与评估的生命周期

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

Fazhong Liu, Zhuoyan Chen, Haozhen Tan, Yan Meng, Guoxing Chen, Haojin Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

本综述针对基于世界模型的具身智能,梳理其生命周期内的各类安全威胁,提出分类法与评估协议,并从多维度构建防御体系,同时指出世界模型兼具安全护盾与安全错觉的双重属性。

中文摘要 AI 辅助

世界模型为具身智能提供预测核心:它们将观测压缩为状态,模拟动作条件下的未来,并支持超越反应式控制的规划。然而,这一预测层开启了新的安全边界——入侵可从数据、传感器、提示或反馈传播至物理动作。本综述未将世界模型视为孤立组件,而是追踪其整个生命周期中的威胁:从数据构建与表示学习,经状态接地与想象,到轨迹评估、执行及通过记忆与工具进行的长期适应。研究表明,熟悉的攻击类型(投毒、后门、对抗样本、传感器欺骗、提示注入、轨迹操纵、供应链攻击)在破坏世界状态、学习动力学、可供性估计或安全成本时具有不同含义。研究还强调了一种双重性:世界模型可作为运行时安全护盾,但当被入侵或过度信任时,会产生预测性安全错觉。本综述提供了生命周期分类法,将现有攻击映射至世界模型安全属性,概述了安全故障的评估协议,并从来源、鲁棒接地、不确定性感知预测、轨迹门控、反馈审计及部署保证等方面构建防御体系。

英文摘要

World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reactive control. This predictive layer, however, opens a new security boundary-compromise can propagate from data, sensors, prompts, or feedback into physical action. Rather than treating world models as an isolated component, this survey traces threats across their entire lifecycle-from data construction and representation learning, through state grounding and imagination, to trajectory evaluation, execution, and long-term adaptation via memory and tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states, learned dynamics, affordance estimates, or safety costs. We also highlight a duality: world models can serve as runtime safety shields, yet when compromised or over-trusted they generate predictive safety illusions. The survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.

↑