arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00267cs.CRcs.AI

无信任的委托:多智能体大语言模型系统中身份、授权与运行时治理的实证差距分析

Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems

Panduranga Sai Varma Dantuluri, Jyotirmoy Sundi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多智能体LLM系统的身份、授权与运行时治理差距,提出不可信模型假设下的安全要求,实现并评估了可缩小差距的授权代理,其能抵御多种攻击且性能开销可忽略。

中文摘要 AI 辅助

自主大语言模型(LLM)智能体越来越多地代表用户行动:它们持有凭证、调用工具和服务,并生成进一步代表其行动的子智能体。这将一个长期存在的分布式系统问题——谁有权限做什么、基于谁的授权——转变为一个紧迫且基本未解决的问题,因为驱动每个智能体的组件是攻击者可以劫持的语言模型。我们认为,智能体安全性必须在不可信模型假设下进行评估:正确的系统应是,即使是完全提示注入的智能体也无法超出明确委托给它的权限。针对这一标准,我们做出三项贡献。第一,我们提出以四个攻击者为核心的多智能体委托威胁模型——混淆代理人、令牌盗窃与重放、提示注入权限提升以及被攻陷的子智能体——并推导了受治理智能体系统必须满足的八项安全要求。第二,我们证明该差距确实存在:模拟常见实践的默认智能体运行时(宽泛的持票凭证、模型内部的授权)无法抵御所有四种威胁,且在四个广泛使用的框架——LangGraph、CrewAI、AutoGen 以及模型上下文协议(MCP)授权模型——中,三个框架未提供内置约束,一个仅提供部分约束;没有任何现有标准单独覆盖该要求集。第三,我们实现并对抗性评估了一个可缩小该差距的授权代理。它可阻止所有四种威胁;它能抵御对其设计的11次直接攻击,且20万枚伪造令牌均未被接受;它能将被攻陷的子智能体限制在其委托任务内(在2000个随机场景中,可访问动作的平均值为1.5,而持票委托下为全部8100个);且它以微秒级成本执行(每次决策约2.6微秒),相对于模型推理可忽略不计。这些原则也在 VotalAI 的 LLM Shield 产品中得到实现。

英文摘要

Autonomous LLM agents increasingly act on a user's behalf: they hold credentials, call tools and services, and spawn sub-agents that act further on their behalf. This turns a long-standing distributed-systems question -- who is authorized to do what, on whose authority -- into an urgent and largely unsolved problem, because the component driving each agent is a language model an adversary can hijack. We argue that agent security must be evaluated under an untrusted-model assumption: a correct system is one in which a fully prompt-injected agent still cannot exceed the authority explicitly delegated to it. Against this standard we make three contributions. First, we give a threat model for multi-agent delegation centered on four adversaries -- confused deputy, token theft and replay, prompt-injection privilege escalation, and compromised sub-agents -- and derive eight security requirements a governed agent system must meet. Second, we show the gap is real: a default agent runtime modeling common practice (broad bearer credentials, authorization gated inside the model) fails all four threats, and across four widely used frameworks -- LangGraph, CrewAI, AutoGen, and the Model Context Protocol (MCP) authorization model -- three provide no built-in confinement and one only partial; no existing standard alone covers the requirement set. Third, we implement and adversarially evaluate an authorization broker that closes the gap. It blocks all four threats; it resists 11 direct attacks on its design and accepts 0 of 200,000 forged tokens; it confines a compromised sub-agent to its delegated task (a mean of 1.5 reachable actions versus all 8,100 under bearer delegation, across 2,000 randomized scenarios); and it enforces at microsecond cost (about 2.6 microseconds per decision), negligible against model inference. These principles are also realized in production in VotalAI's LLM Shield.

发表机构

  • VotalAI(沃塔尔人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑