有界智能体:多智能体AI系统的委托安全
Bounded Agents: Delegation Security for Multi-Agent AI Systems
浏览论文内容
中文总结 AI 辅助
该研究针对多智能体AI系统的委托安全问题,提出智能体委托链(APC)架构,经实验可有效阻止提示注入攻击,降低数据窃取、破坏及操纵风险,授权延迟低且效用损失可控。
中文摘要 AI 辅助
基于大语言模型(LLM)的智能体可代表用户执行访问云服务、调用工具或激活其他智能体等操作。会话开始时,智能体的权限已设定且保持静态,每个请求会被单独评估,不考虑之前的操作。在权限范围内,智能体可能会违背委托任务行事、将各自允许的操作组合成被禁止的结果,或在未加限制的情况下将权限委托给子智能体。提示注入仅在智能体有权执行此类操作时才会构成风险,因此这是授权架构的问题,而非仅模型的问题。智能体委托链(APC)用于跟踪从一个委托方到下一个委托方的委托权限,它会基于累积的会话状态,通过六项授权检查来评估每个请求。APC会传递并限制委托的范围和预算,利用组合闭包,APC会根据之前的操作检查请求,以防止被禁止的组合,并在模型之外执行决策。我们证明了APC实现的爆炸半径单调性和组合正确性,其中组合正确性仅限于在完整限制集和序列化准入下的被禁止组合。我们评估了3154个实例,包括InjecAgent、AgentDojo和ASB。我们的受损模型评估通过在第一次合法工具调用后插入真实攻击调用,独立于模型行为测试APC。AgentDojo的渗出率在所有四个领域从75%-100%降至0%;APC阻止了全部544起InjecAgent数据窃取案例。意图绑定将破坏率从38.6%降至4.0%,操纵率从90.5%降至12.1%。授权延迟在空闲主机上的第99百分位为0.24毫秒;在949个AgentDojo任务注入对中,两种设置下的效用分别降低了8.6和13.9个百分点。实现、评估工具和数据均公开可用。
英文摘要
LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static, and each request is evaluated independently, without considering prior actions. Within its permissions, an agent may act contrary to the delegated task, combine individually permitted actions into a prohibited outcome, or delegate authority to a sub-agent without limiting it. A prompt injection poses a risk only if the agent has authority to perform such actions; this is therefore a problem of authorization architecture, not just the model. The Agentic Principal Chain (APC) tracks delegated authority from one principal to the next. APC evaluates each request against the accumulated session state using six authorization checks. APC carries forward and restricts delegated scope and budgets. Using composition closure, APC checks requests against prior actions to prevent prohibited combinations and enforces the decision outside the model. We prove Blast Radius Monotonicity and Composition Soundness for APC implementations; Composition Soundness is limited to prohibited combinations under a complete restriction set and serialized admission. We evaluated 3,154 instances including InjecAgent, AgentDojo, and ASB. Our compromised-model evaluation tests APC independently of model behavior by inserting the ground-truth attack call after the first legitimate tool call. AgentDojo exfiltration fell from 75-100% to 0% across all four domains; APC blocked all 544 InjecAgent data-stealing cases. Intent binding reduced destruction from 38.6% to 4.0% and manipulation from 90.5% to 12.1%. Authorization latency was 0.24 ms at the 99th percentile on an idle host; across 949 AgentDojo task-injection pairs, utility was 8.6 and 13.9 percentage points lower in the two settings. Implementation, evaluation tools, and data are publicly available.