面向安全的大语言模型智能体:规范、验证与执行的综述
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
- The University of Manchester(曼彻斯特大学)
- The University of Greenwich(格林威治大学)
- Technology Innovation Institute(技术创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该综述针对LLM智能体缺乏形式化任务级安全保障的问题,通过系统分析38项研究,揭示规范瓶颈等四项关键发现,提出三级分类体系与十大问题研究议程,为可信智能体AI研究提供方向。
AI中文摘要:
大语言模型(LLM)智能体越来越多地执行不可逆的现实世界操作,包括数据库更新、API调用、文件操作以及工具的自主使用。然而,现有系统均未为这些智能体生成的计划提供基于形式化的任务级安全保障。相关研究在规范、验证和执行三个方向上分散,限制了对现有方法优缺点的理解。为填补这一空白,我们开展了PRISMA 2020系统综述,检索了六个学术数据库中2022年至2026年发表的38项研究。我们的分析揭示了四项关键发现:第一,规范瓶颈仍是主要挑战:自然语言到形式化的翻译仅达到24%至35%的语义正确性,损害了下游验证效果;第二,运行时监控是最成熟的执行策略,在受控场景中可减少40%至65%的不安全操作,但无法提供完整的安全保障;第三,验证器代价表明,即使阻断94%的不安全操作,仍可能导致安全任务完成率不足5%,因为智能体会利用替代的不安全路径;第四,现有方法无法同时实现可靠性、可扩展性、语义正确性和任务级安全保持。我们贡献了一个三级分类体系、现有技术的比较分析、关于验证器代价的证据综合,以及一份针对可信智能体AI的十大问题研究议程。
英文摘要:
LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate. Research remains fragmented across specification, verification, and enforcement, limiting understanding of the strengths and limitations of existing approaches. To address this gap, we conducted a PRISMA 2020 systematic review of 38 studies published between 2022 and 2026 and retrieved from six academic databases. Our analysis reveals four key findings. First, the specification bottleneck remains the primary challenge: natural-language-to-formal translation achieves only 24% to 35% semantic correctness, undermining downstream verification. Second, runtime monitoring is the most mature enforcement strategy, reducing unsafe actions by 40% to 65% in controlled settings, but it does not provide complete safety guarantees. Third, the verifier tax shows that blocking 94% of unsafe actions can still result in less than 5% safe task completion because agents exploit alternative unsafe paths. Finally, no existing approach simultaneously achieves soundness, scalability, semantic correctness, and task-level safety preservation. We contribute a three-level taxonomy, a comparative analysis of existing techniques, a synthesis of evidence on the verifier tax, and a ten-problem research agenda for trustworthy agentic AI.