关于智能体大语言模型漏洞的理解、识别与缓解
On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究针对智能体大语言模型安全开展系统文献综述,提出四层漏洞分类法,发现攻击研究多于防御研究、感知层漏洞研究占比偏高等现状,识别7个开放问题及架构耦合的核心不安全根源。
中文摘要 AI 辅助
大语言模型(LLMs)已从无状态的对话界面转变为具备多步骤规划、工具调用、代码执行及持久记忆能力的自主智能体。当这些智能体拥有调用API、修改文件、查询数据库等真实世界权限时,推理步骤被破坏可能引发未授权数据访问、不可逆状态变更或级联故障,但安全研究领域尚未跟上该发展节奏。为量化该领域现状,我们依据PRISMA 2020指南在6个数据库开展系统文献综述,筛选743条记录后保留85篇2023至2025年关于智能体大语言模型安全的论文。攻击研究与防御研究的比例为3.9:1;感知层漏洞(提示注入、越狱、对抗扰动)占比最高,达66%,而动作层漏洞(工具滥用、代码注入、沙箱逃逸)仅占4.7%,与真实世界风险不匹配;代码执行安全占3.5%,工具增强智能体相关研究占12%。我们提出涵盖感知层、大脑层、动作层、交互层的四层分类法,映射13种漏洞类型,并识别出以漏洞遏制为核心的7个开放问题。智能体大语言模型的不安全源于架构耦合,薄弱的隔离机制使漏洞可跨层传播。
英文摘要
Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agents operate with real-world privileges---calling APIs, modifying files, and querying databases---a compromised reasoning step can trigger unauthorized data access, irreversible state changes, or cascading failures, yet the security research community has not kept pace. To quantify the state of the field, we conducted a systematic literature review under PRISMA 2020 guidelines across six databases, screening 743 records and retaining 85 papers (2023--2025) on agentic LLM security. Attack research outpaces defense work by 3.9:1. Perception-layer vulnerabilities (prompt injection, jailbreaking, adversarial perturbations) dominate, accounting for 66\% of papers, while action-layer vulnerabilities (tool misuse, code injection, sandbox escape) appear in only 4.7\%, misaligned with real-world risk. Code execution security accounts for 3.5\%, and tool-augmented agents 12\%. We contribute a four-layer taxonomy mapping 13 vulnerability types across perception, brain, action, and interaction layers, and identify seven open problems centered on containment. Agentic LLM insecurity stems from architectural coupling, where weak isolation allows vulnerabilities to propagate across layers.
发表机构
- Florida International University(佛罗里达国际大学)
- Middle Tennessee State University(中田纳西州立大学)
- New Jersey Institute of Technology(新泽西理工学院)
机构由 AI 辅助整理,请以论文原文为准。