arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16843cs.RO

基于基础模型的具身智能体的安全性:攻击面、攻击、防御与评估

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究以信任边界为核心,针对基于基础模型的具身智能体安全,划分了五个层级与十二个攻击面,分析了58种攻击、61种防御的现状,指出部分领域研究不足并提出开放挑战。

中文摘要 AI 辅助

基础模型正越来越多地用于具身智能体的感知、推理、规划和动作生成,由此产生的安全风险可从数字输入传播至物理行为。现有调查常通过越狱、提示注入、后门、投毒或对抗样本等机制来组织威胁,但这些类别无法一致地识别攻击者首次侵入具身控制循环的位置。我们开展了一项以信任边界为核心的基于基础模型的具身智能体安全性调查,采用「首次受侵信任边界」原则,将攻击面与攻击机制分离,并将系统划分为五个层级和十二个攻击面,涵盖模型供应链、用户指令、上下文与记忆、物理语义环境、多模态感知、世界状态、内部推理、任务规划、动作接口、中间件、多智能体通信及执行控制。基于截至2026年8月15日收集的58条攻击记录和61条防御记录,我们分析了代表性攻击、跨层级传播、防御部署位置及评估实践。定量分析显示,攻击研究集中于多模态感知和动作接口,而防御尤其集中于动作级和运行时保护;上下文与长期记忆、中间件与网络、世界状态完整性及多智能体信任则相对未得到充分探索。我们最后总结了状态溯源、组合防御、长时程攻击传播、物理可实现性、拜占庭多机器人行为及统一闭环评估等方面的开放挑战。

英文摘要

Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks, prompt injection, backdoors, poisoning, or adversarial examples, but these categories do not consistently identify where an adversary first enters the embodied control loop. We present a trust-boundary-centric survey of foundation-model-powered embodied-agent security. Using a first-compromised-trust-boundary principle, we separate attack surface from attack mechanism and organize the system into five layers and twelve attack surfaces spanning the model supply chain, user instructions, context and memory, physical semantic environments, multimodal perception, world state, internal reasoning, task planning, action interfaces, middleware, multi-agent communication, and execution control. Based on 58 attack records and 61 defense records collected through August 15, 2026, we analyze representative attacks, cross-layer propagation, defense placement, and evaluation practices. Our quantitative analysis shows that attack research is concentrated on multimodal perception and action interfaces, while defenses are especially concentrated on action-level and runtime protection. Context and long-term memory, middleware and networking, world-state integrity, and multi-agent trust remain comparatively underexplored. We conclude with open challenges in state provenance, compositional defenses, long-horizon attack propagation, physical realizability, Byzantine multi-robot behavior, and unified closed-loop evaluation.

发表机构

  • Wuhan University(武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

↑