arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隔离作为大语言模型智能体系统安全的头等原则:概念、分类法、挑战及未来方向

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Sirui Zhang, Hanyu Yang, Changxuan Fan, Zhongwei Xie, Hongyu Luo, Wun Yu Chan, Wei Fan, Haoran Li, Yangqiu Song

arXiv 2607.12406首次发表:更新:

发表机构

HKUST; NYU; SWUPL(香港科技大学; 纽约大学; 西南政法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

探讨大语言模型智能体系统安全问题,将隔离作为头等原则,用边界中心分类法组织文献,助于识别隔离丧失位置、妥协传播方式及相关防御,还总结了故障路径,讨论挑战并勾勒研究议程。

AI 中文摘要

大语言模型智能体作为系统“大脑”的能力,从根本上扩展了分析范围,超越了独立模型。因此,安全不仅关乎输入输出内容对齐,还涉及系统行为和现实世界执行结果。然而,当前文献在攻击类型、应用和基准方面较为分散。本文将隔离视为大语言模型智能体系统安全的头等原则,通过用户输入、工具访问、执行通道、智能体间通信和环境源上下文的分离来实现。我们用一种以边界为中心的分类法组织文献,涵盖五个边界:用户-智能体、智能体-工具、智能体-执行、智能体-智能体和系统-环境。这有助于识别隔离丧失首先发生的位置、妥协如何跨边界传播以及每个接口最相关的防御措施。我们还总结了跨边界故障路径,讨论了开放挑战,并概述了未来智能体系统中通过构建实现隔离的研究议程。

英文摘要

The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behavior and real-world execution outcomes. However, the current literature is fragmented across attack types, applications, and benchmarks. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share the same structural cause, and how they spread through an agent workflow. In this survey, we treat isolation as a first-class principle for LLM-agent system safety. By isolation, we refer to the separation of user inputs, tool access, execution channels, inter-agent communication, and environment-originated context. We organize the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment. This view helps identify where the loss of isolation first occurs, how compromise propagates across boundaries, and which defenses are most relevant at each interface. We also summarize cross-boundary failure paths, discuss open challenges, and outline a research agenda for isolation-by-construction in future agent systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑