遏制自主操作者:在Kubernetes上保护AI智能体的纵深防御框架与参考架构
Containing the Autonomous Operator: A Defense-in-Depth Framework and Reference Architecture for Securing AI Agents on Kubernetes
浏览论文内容
中文总结 AI 辅助
本文针对Kubernetes上LLM智能体面临提示注入威胁,提出威胁模型、九项设计原则及七层纵深防御框架,并给出参考架构与定性评估,强调安全边界在基础设施而非模型。
中文摘要 AI 辅助
大型语言模型(LLM)智能体正从聊天界面走向基础设施运维领域,它们读取遥测数据、调用工具、生成并执行代码,以及改变生产环境Kubernetes集群的状态。这打破了传统云原生安全所依赖的边界:数据与控制之间的界限。智能体仅读取的内容(一行日志、一张工单、一个工具描述)就能改变其行为。本文主张,模型不应被视为安全边界,因此Kubernetes上的智能体安全本质上是一个基础设施问题:所有保证必须在假设智能体已被提示注入完全攻破的情况下依然成立。我们贡献了:(i)一个威胁模型和针对在Kubernetes上及内部运行的智能体的十类威胁分类法,与新兴的OWASP智能体应用指南保持一致;(ii)九项设计原则,核心是在工具边界实现完全中介,并打破不可信输入、敏感访问和外部出口的组合;(iii)一个七层纵深防御框架,将每项原则映射到原生或广泛采用的Kubernetes机制:工作负载身份、RBAC和ValidatingAdmissionPolicy、通过SIG Apps Agent Sandbox项目实现的gVisor/Kata沙箱、FQDN感知的出口策略、带有工具参数策略即代码的智能体/MCP网关,以及eBPF运行时强制;(iv)一个参考架构,包含具体的策略工件和针对Amazon EKS、Azure Kubernetes Service和Google Kubernetes Engine的逐层绑定;(v)一个定性评估,包括威胁控制覆盖矩阵和四次攻击演练,并提出了经验性方法论。我们未报告实测的攻击成功率或开销数据;相反,我们识别了残余风险以及验证该框架所需的测量方法。
英文摘要
Large language model (LLM) agents are moving from chat interfaces into infrastructure operations, where they read telemetry, call tools, generate and execute code, and change the state of production Kubernetes clusters. This collapses a boundary that conventional cloud-native security assumes: the boundary between data and control. Content that an agent merely reads (a log line, a ticket, a tool description) can redirect what it does. This paper argues that the model must not be treated as a security boundary and that agent safety on Kubernetes is therefore an infrastructure problem: every guarantee must continue to hold under the assumption that the agent is fully compromised by prompt injection. We contribute (i) a threat model and ten-class threat taxonomy for agents operating on and within Kubernetes, aligned with emerging OWASP guidance for agentic applications; (ii) nine design principles, centered on complete mediation at the tool boundary and on breaking the combination of untrusted input, sensitive access, and external egress; (iii) a seven-layer defense-in-depth framework that maps each principle to native or widely adopted Kubernetes mechanisms: workload identity, RBAC and ValidatingAdmissionPolicy, gVisor/Kata sandboxing via the SIG Apps Agent Sandbox project, FQDN-aware egress policy, an agent/MCP gateway with policy-as-code over tool arguments, and eBPF runtime enforcement; (iv) a reference architecture with concrete policy artifacts and per-layer bindings for Amazon EKS, Azure Kubernetes Service, and Google Kubernetes Engine; and (v) a qualitative evaluation comprising a threat-control coverage matrix and four attack walkthroughs, with a proposed empirical methodology. We report no measured attack-success or overhead figures; instead we identify residual risks and the measurements needed to validate the framework.