AI 中文总结
针对多智能体AI系统的安全风险,本文提出Agent Containment架构,通过六项约束实施「提议-验证-行动-验证」模型,可防御多种威胁,案例验证其能产生合规可审计结果。
AI 中文摘要
多智能体AI系统正越来越多地部署在自主协作、工具使用和持续学习会引入新型安全与治理风险的场景中。针对多智能体系统安全的 containment 方法仍未充分发展,主要依赖设计后、有时甚至是实现与部署后的附加层。本文提出一种 Agent Containment 架构,将安全视为通过一组显式约束实施的架构属性,这些约束限定了多智能体系统的设计空间。该架构引入了一种新颖的形式化组织映射,将标准系统分析工件映射到所提约束系统下可机器验证的合约。架构包含六个相互作用的约束:职责分配分离、部署前一致性检查、价值流绑定、时间隔离、积累前的严格知识验证,以及结构与流程完整性的确定性验证。这些约束共同实施了「提议-验证-行动-验证」执行模型,其中所有操作都有合约定义、独立验证且可追溯至特定执行上下文。本文提出了将这些约束与针对关键威胁类别(包括提示注入、协调器操纵、跨会话状态投毒及涌现智能体合谋)的防御相关联的命题与正确性推理论证。一份简历筛选案例研究展示了该架构如何在对抗条件下产生可审计、符合策略的结果。该工作明确将结构完整性与语义安全分离,限定了残余风险,同时使残余语义风险明确且可衡量。
英文摘要
Multi-agent AI systems are increasingly deployed in contexts where autonomous coordination, tool use, and continuous learning introduce novel security and governance risks. The containment approach to multi-agent system security remains underdeveloped, primarily resorting to add-on layers post-design and sometimes post-implementation and deployment. This paper proposes an Agent Containment Architecture that treats security as an architectural property enforced through a set of explicit constraints that bound the design space of multi-agent systems. The architecture proposed in this paper introduces a novel formal organizational mapping from standard systems analysis artifacts to machine-verifiable contracts under a proposed constraint system. The architecture introduces six interacting constraints: separation of responsibility assignments, pre-deployment coherence checking, value stream binding, temporal isolation, strict knowledge verification before accumulation, and deterministic verification of structural and process integrity. Together, these constraints enforce a Propose-Verify-Act-Verify execution model in which all operations are contractually defined, independently verified, and traceable to specific execution contexts. The paper presents propositions and correctness reasoning arguments linking these constraints to defenses against key threat classes, including prompt injection, orchestrator manipulation, cross-session state poisoning, and emergent agent collusion. A resume screening case study demonstrates how the architecture produces auditable, policy-compliant outcomes under adversarial conditions. The work explicitly separates structural integrity from semantic safety, bounding residual risks while making residual semantic risk explicit and measurable.
CommentsAccepted to IntelliSys'2026. To Appear in Springer series "Lecture Notes in Networks and Systems" in 2026