多智能体LLM流水线中的信任传播与结构遏制
Trust propagation and structural containment in Multi-agent LLM pipelines
- BRAC University(BRAC大学)
- University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过四智能体LangGraph流水线实验,提出判断绕过率(JBR)指标,证明结构性授权层可遏制被攻破智能体的行为,即使上游LLM判断失败,同时将误报率从49%降至7%。
AI中文摘要:
多智能体LLM系统日益自动化涉及不同权限级别智能体的任务,这造成了一种安全风险:被攻破的低权限智能体可以影响高权限智能体并触发未经授权的操作。我们研究了由Supervisor、Researcher、Validator和Executor组成的四智能体LangGraph流水线中的攻击传播。我们评估了通过检索文档中嵌入的伪造审批进行的共享内存投毒和间接提示注入。我们将Validator的判断与使用任务绑定签名令牌和独立验证的策略预言机的独立授权层进行比较。我们的贡献是对攻击传播的实证研究、授权边界的组件级消融,以及衡量被攻击智能体而非最终行动处受损程度的判断绕过率(JBR)。在三个随机种子和60个标记任务中,内存投毒在每次未防御试验中都达到了执行阶段。启用授权后,它实现了100%的JBR但0%的不安全行动率,表明Validator可以保持受损状态而执行被遏制。针对拥有签名秘密的攻击者,策略预言机提供了观察到的遏制效果,而独立编写的最小权限策略保持了这一结果。Observer层将智能体劫持的误报率从49%降至7%,且不削弱执行级安全性。这些结果表明,即使上游LLM判断失败,结构性授权也能遏制受损智能体的行为。
英文摘要:
Multi-agent LLM systems increasingly automate tasks involving agents with different levels of privilege, creating a security risk in which a compromised low-privilege agent can influence a higher-privilege agent and trigger an unauthorized action. We study attack propagation in a four-agent LangGraph pipeline comprising a Supervisor, Researcher, Validator, and Executor. We evaluate shared-memory poisoning and indirect prompt injection through a forged approval embedded in a retrieved document. We compare the Validator's judgment with an independent authorization layer using task-bound signed tokens and a separately verified policy oracle. Our contribution is an empirical study of attack propagation, a component-level ablation of the authorization boundary, and the Judgment Bypass Rate (JBR), which measures compromise at the attacked agent rather than at the final action. Across three seeds and 60 labeled tasks, memory poisoning reaches execution in every undefended trial. With authorization enabled, it achieves 100% JBR but 0% Unsafe Action Rate, showing that the Validator can remain compromised while execution is contained. Against an attacker possessing the signing secret, the policy oracle provides the observed containment, while an independently authored least privilege policy preserves this result. An Observer layer reduces the false-positive rate for agent hijacking from 49% to 7% without weakening execution-level security. These results show that structural authorization can contain compromised agent behavior even when upstream LLM judgment fails.