发表机构
University of Maryland Baltimore County; Kent State University; Tennessee State University(马里兰大学巴尔的摩郡分校; 肯特州立大学; 田纳西州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多智能体LLM流水线的结构性漏洞,通过在GPT-5-mini等模型上的实验,发现对抗性漏洞源于架构而非模型能力,推动流水线级防御的发展。
AI 中文摘要
多智能体大语言模型(LLM)流水线将多个专用语言模型智能体编排为结构化工作流,中间输出在智能体间传递以解决复杂任务。该设计引入了单智能体场景不存在的安全缺口:一旦某个智能体接受对抗性内容,它就会作为可信输入在整个流水线中传播。我们认为该漏洞源于缺乏边界验证,这是一种安全原语,用于在数据跨智能体边界时(包括内容、身份、执行意图和状态完整性)强制进行显式验证。没有此类验证,现代流水线嵌入了对抗性不稳健的隐式信任假设,从而产生结构上不同的攻击面(例如内容注入、智能体冒充、计划偏离和内存投毒)。利用来自GAIA和SWE-Bench基准的带注释生产轨迹,我们表明这些漏洞在良性部署中出现,且在很大程度上避开了现有评估框架。我们进一步在受控多智能体场景中实现这些失效模式,并在相同流水线配置下对GPT-5-mini、Claude Sonnet 4.5和Kimi K2.5进行评估。结果显示,攻击成功与流水线结构而非模型能力一致,表明对抗性漏洞本质上是一种架构属性,推动转向流水线级防御。
英文摘要
Multi-agent LLM pipelines orchestrate multiple specialized language model agents into structured workflows where intermediate outputs are passed across agents to solve complex tasks. This design introduces a security gap absent in single-agent settings: once an agent accepts adversarial content, it is propagated as trusted input throughout the pipeline. We argue that this vulnerability stems from the absence of boundary verification, a security primitive that enforces explicit validation of data as it crosses inter-agent boundaries, including content, identity, execution intent, and state integrity. Without such verification, modern pipelines embed implicit trust assumptions that are not adversarially robust, giving rise to structurally distinct attack surfaces (e.g., content injection, agent impersonation, plan deviation, and memory poisoning). Leveraging annotated production traces from the GAIA and SWE-Bench benchmark, we show that these vulnerabilities arise in benign deployments and largely evade existing evaluation frameworks. We further operationalize these failure modes within a controlled multi-agent setting and evaluate them across GPT-5-mini, Claude Sonnet 4.5, and Kimi K2.5 under identical pipeline configurations. The results reveal that attack success aligns with pipeline structure rather than model capability, indicating that adversarial vulnerability is fundamentally an architectural property and motivating a shift toward pipeline-level defenses.
CommentsThis paper has been accepted at the 2026 IEEE Global Communications Conference (GLOBECOM)