发表机构
Florida International University(佛罗里达国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究系统化梳理智能体AI时代的越狱攻击与防御,发现原生对齐不等于对抗鲁棒性,防御依赖具体情境且代价高,低最终成功率可能掩盖中间环节风险,主张转向跨层执行感知的安全防护。
AI 中文摘要
大型语言模型(LLM)正迅速从对话助手演变为智能体AI系统,这些系统能够进行推理、规划、调用工具、维护持久记忆、与其他智能体通信,并执行多步骤任务。与此同时,现代模型展现出比早期几代模型(许多越狱攻击和防御最初正是在这些早期模型上研究的)显著更强的原生安全对齐能力。这一转变引发了一个根本性问题:\textit{在当代LLM和智能体AI时代,哪些已确立的越狱安全发现仍然有效?}我们通过一项系统化知识(SoK)研究来回答这一问题,该研究围绕完整的智能体执行流水线重新构建了越狱安全框架。我们开发了涵盖用户交互、规划与推理、记忆、工具使用以及智能体间通信的攻击与防御统一分类法,并引入了一个安全-效用-效率评估框架,该框架将原生有害提示安全性、对抗性越狱鲁棒性和智能体级安全结果分离开来。我们进一步在共同的智能体框架内对代表性攻击和防御进行了受控实证研究。我们的结果揭示了三个重要差距。第一,强大的原生对齐并不意味对对抗性越狱具有鲁棒性。第二,防御有效性高度依赖于模型、攻击和组件,并且可能以过度拒绝、效用和延迟方面的巨大成本为代价。第三,低最终响应攻击成功率可能掩盖严重的中间环节妥协:即使最终响应被成功过滤,规划、记忆和工具交互仍可能不安全。这些发现促使我们从以响应为中心的越狱防御转向跨层、执行感知的安全防护,在保护智能体状态、组件转换和外部操作的同时,保持实际效用和效率。
英文摘要
Large language models (LLMs) are rapidly evolving from conversational assistants into agentic AI systems that reason, plan, invoke tools, maintain persistent memory, communicate with other agents, and execute multi-step tasks. At the same time, modern models exhibit substantially stronger native safety alignment than earlier generations on which many jailbreak attacks and defenses were originally studied. This shift raises a fundamental question: \textit{which established jailbreak-security findings remain valid in the era of modern LLMs and agentic AI?} We address this question through a Systematization of Knowledge (SoK) that reframes jailbreak security around the full agentic execution pipeline. We develop unified taxonomies of attacks and defenses spanning user interaction, planning and reasoning, memory, tool use, and inter-agent communication, and introduce a security--utility--efficiency evaluation framework that separates native harmful-prompt safety, adversarial jailbreak robustness, and agent-level security outcomes. We further conduct a controlled empirical study of representative attacks and defenses within a common agentic framework. Our results reveal three important gaps. First, strong native alignment does not imply robustness to adversarial jailbreaks. Second, defense effectiveness is highly model-, attack-, and component-dependent and can come at substantial cost in over-refusal, utility, and latency. Third, low final-response attack success can mask severe intermediate compromise: planning, memory, and tool interactions may remain unsafe even when the final response is successfully filtered. These findings motivate a shift from response-centric jailbreak defense toward cross-layer, execution-aware security that protects agent state, component transitions, and external actions while preserving practical utility and efficiency.