发表机构
The Hong Kong University of Science and Technology; Renmin University of China(香港科技大学; 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对具身智能体的越狱攻击,本文首次系统评测六种防护栏,构建可插拔框架,从防御有效性、可用性和效率三维度评估,发现无单一防护栏全面占优,并提供选择与设计指导。
AI 中文摘要
由大语言模型和视觉语言模型驱动的具身智能体正越来越多地部署在物理环境中,但越狱攻击可能诱使这些智能体执行物理上有害的行为。越来越多的防护栏方法被提出,以在危险行为执行前进行拦截,然而现有的安全基准评测的是具身模型本身,这使得这些防护栏在实践中究竟能在多大程度上防御具身智能体仍不明确。我们首次对具身智能体的越狱防护栏进行了系统性评测。为了在相同条件下比较防护栏,我们构建了一个可插拔的评测框架,将具身智能体视为固定后端,并将每个防护栏视为可在感知、规划或控制阶段进行干预的模块。我们让六种代表性防护栏面对基于模板和自动化的越狱攻击以及安全指令,并在系统层面从三个维度对其进行评估:防御有效性,通过模拟器中的绕过率和危险成功率来衡量;可用性,通过安全指令上的误报率和任务完成率来衡量;以及效率,通过运行时增加的延迟开销来衡量。对跨越不同干预阶段、决策机制和输入模态的防护栏进行的实验揭示了这三个维度之间的明显权衡,并表明没有任何单一防护栏在所有设置中占据主导地位。我们进一步分析了干预阶段、决策机制和输入模态如何影响安全结果,并为选择和设计具身智能体的防护栏提供了实用指导。
英文摘要
Embodied agents powered by large language models and vision-language models are increasingly deployed in physical environments, but jailbreak attacks can induce these agents to perform physically harmful actions. A growing number of guardrail methods have been proposed to intercept dangerous behavior before it is executed, yet existing safety benchmarks evaluate the embodied models themselves, leaving it unclear how well these guardrails actually defend an embodied agent in practice. We present the first systematic evaluation of jailbreak guardrails for embodied agents. To compare guardrails under identical conditions, we build a pluggable evaluation framework that treats the embodied agent as a fixed backend and each guardrail as a module that can intervene at the perception, planning, or control stage. We subject six representative guardrails to template-based and automated jailbreak attacks as well as safe instructions, and assess them at the system level along three dimensions: defense effectiveness, measured by the bypass rate and the hazard success rate in the simulator; usability, measured by the false-positive rate and the task completion rate on safe instructions; and efficiency, measured by the latency overhead added at runtime. Experiments on guardrails that span different intervention stages, decision mechanisms, and input modalities reveal a clear trade-off among the three dimensions, and show that no single guardrail dominates in all settings. We further analyze how intervention stage, decision mechanism, and input modality shape safety outcomes, and we offer practical guidance for selecting and designing guardrails for embodied agents.