发表机构
Accenture Labs(埃森哲实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自主AI代理易受提示注入和凭证泄露的问题,本文提出一种基于能力安全与去中心化身份的令牌框架,通过密码学绑定令牌和范围衰减验证,在模型外部实现安全边界。
AI 中文摘要
完全自主的交易代理的前景直到高能力语言模型的出现才出现在地平线上。借助此类模型,自适应任务编排和独立(但受约束)决策的操作优势对企业及个人而言极具吸引力。然而,每个此类代理都携带着一个严重的攻击面,即提示注入,它可能破坏上下文中可能已设置的任何软性“护栏”。这些攻击的后果包括凭证外泄,如果得不到缓解,将因信任破坏而使整个代理类别无法使用。此外,考虑到自主代理的增长预期,为其构建安全框架需要某种形式的去中心化以实现扩展。借鉴拟议的OAuth代理授权配置文件以及W3C DID和VC标准,我们提出一个基于能力安全原则和去中心化代理身份的框架,其中代理基于令牌访问服务,这些令牌在密码学上绑定到代理和签发者的身份,并指定其范围。服务可以在对任何给定令牌采取行动之前,验证委托链仅涉及范围衰减。我们表明,这样一个位于语言模型上下文窗口之外的安全模块中的层,能够使代理在可执行的安全边界内行动。
英文摘要
The prospect of fully autonomous transactional agents did not appear on the horizon until the advent of high capability language models. With such models, the operational benefits of adaptive task orchestration and independent (but constrained) decision making are tantalizing for enterprises and individuals alike. However, each such agent carries with it a serious attack surface in the form of prompt injection which can compromise any soft "guard rails" that may have been placed in context. The consequences of these attacks include credential ex-filtration which, if left unmitigated, renders the whole category of such agents unusable due to breach of trust. Furthermore, expecting a growth of autonomous agents, a security framework for them would require a form of decentralization to scale. Drawing on the proposed OAuth Agent Authorization Profile and the W3C DID and VC standards, we propose a framework based on the principles of capability based security with decentralized agent identity whereby agents access services based on tokens that are cryptographically bound to the agent's and issuer's identities and specify their scope. Services can validate that the delegation chain only involves scope attenuation before acting on any given token. We show that such a layer that lives outside the language model's context window in a secure module can enable agents to act within enforceable security boundaries.
Comments9 pages, 9 figures, 3 tables