发表机构
Westcliff University; University of the Cumberlands; Delta Air Lines, Inc.; Georgia Institute of Technology(韦斯特克利夫大学; 坎伯兰大学; 达美航空股份有限公司; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对工具使用型AI智能体的授权安全问题,提出主体层级框架和四层参考架构,并识别运行时执行与聚合边界为关键未解空白。
AI 中文摘要
工具使用型人工智能(AI)智能体——即自主调用应用程序编程接口(API)、数据库、浏览器以及诸如模型上下文协议(MCP)等智能体间协议的系统——正成为生产基础设施。然而,关于智能体何时被授权代表人类行事的的安全模型仍不完善。可信赖的人机AI系统要求每个重要的智能体行为都能追溯到人类主体,受限于该人类实际委托的范围,并且事后可提出异议;很少有文档记载的部署能够端到端地可靠满足全部三项属性。现有文献孤立地处理该问题的碎片,例如非人类身份的凭证管理、经典访问控制模型、提示注入和审计追踪,却很少关注授权决策点本身——即工具调用发生的时刻——以及使该决策正确、可执行且可问责的机制。本综述引入了一个涵盖人类用户、操作员/部署者、编排智能体、子智能体和工具端点的主体层级作为组织框架,并考察了五个相互依赖的层:智能体身份与凭证生命周期;跨多跳链的委托与范围传播;在策略执行点(PEP)的运行时执行与即时授权;作为破坏主体层级的授权绕过的提示注入;以及可审计性、来源和不可否认性。基于对从约180个候选来源中筛选出的、发表于2023年至2026年间的89份主要文献的结构化叙述性综述,我们提出了七项结构性要求,推导出一个四层参考架构,将要求应用于三种可部署的参考配置,并确定运行时执行和聚合边界是主要未解决的空白。
英文摘要
Tool-using artificial intelligence (AI) agents, systems that autonomously invoke application programming interfaces (APIs), databases, browsers, and inter-agent protocols such as the Model Context Protocol (MCP), are becoming production infrastructure. Yet the security model governing when an agent is authorized to act on a human's behalf remains underdeveloped. Trustworthy human-AI systems require that every consequential agent action be traceable to a human principal, bounded by what that human actually delegated, and contestable after the fact; few documented deployments satisfy all three properties reliably and end to end. Existing literature addresses fragments of this problem in isolation, credential management for non-human identities, classical access control models, prompt injection, and audit trails, while giving little attention to the authorization decision point itself, the moment a tool invocation occurs, and mechanisms that make that decision correct, enforceable, and accountable. This review introduces a principal hierarchy spanning human user, operator/deployer, orchestrator agent, sub-agent, and tool endpoint as an organizing framework, and examines five interdependent layers: agent identity and credential lifecycle; delegation and scope propagation across multi-hop chains; runtime enforcement and just-in-time authorization at policy enforcement points (PEPs); prompt injection as an authorization bypass that breaks the principal hierarchy; and auditability, provenance, and non-repudiation. Drawing on a structured narrative review of 89 primary sources screened from approximately 180 candidates published between 2023 and 2026, we propose seven structural requirements, derive a four-layer reference architecture, apply the requirements to three deployable reference configurations, and identify runtime enforcement and aggregation bounds as the principal unresolved gaps.
Comments70 pages, 8 figures, 10 tables