从单体到群体:多智能体Web系统转型中的攻击面演化研究
From Monoliths to Swarms: A Study of Attack Surface Evolution in the Transition to Multi-Agent Web Systems
浏览论文内容
中文总结 AI 辅助
本文针对从单智能体到多智能体Web系统转型的安全问题,提出MAS攻击向量分类法,构建WebMASLab测试平台,发现新型电话环路攻击对多数前沿多智能体模型威胁显著,提示防御通用性有限,揭示多智能体架构的新型安全风险。
中文摘要 AI 辅助
基于大语言模型(LLM)的Web智能体正日益从单智能体系统(SAS)向多智能体系统(MAS)演进。尽管MAS可通过将复杂任务分解给专门的子智能体来提升任务性能,但这种角色分解会引入SAS中不存在的新型结构性攻击面,且这一扩展的攻击面仍未被充分理解和归类。为解决该问题,本文提出一种分类法,用于归类基于Web的MAS特有的攻击向量,涵盖因多智能体参与而引入或放大的漏洞。我们还提出了一个测试平台WebMASLab,用于针对完全外部、仅基于Web的敌手分析Web智能体的安全性。为隔离架构的影响,我们保持用户任务、工具面和浏览器底层固定,对比单智能体与多智能体设置。我们在三种条件(基线、提示加固、启用推理)下评估了三种对抗场景,其中包括一种新型MAS特有的电话环路攻击,该攻击利用跨智能体委托创建循环任务环路。该攻击对SAS无效,但在评估的四个前沿模型中的三个(Claude Sonnet 4.5、GPT-5.2、GPT-5.4)作用下会破坏MAS,基线下平均成功率为80%。仅第四个模型Claude Sonnet 4.6以92%的检测率抵御了该攻击,其余模型基线下检测率为0%,其中一个模型经提示加固后检测率达到33%。我们还表明,明显的防御措施无法通用:提示加固使一个模型的攻击成功率(ASR)从100%降至8%,而对其他模型仅实现了有限降低。我们的研究结果表明,从单智能体Web系统向多智能体Web系统的转变改变了安全格局,角色专业化不仅可能带来性能优化,还会引入需要进一步研究和防御的新型架构风险。
英文摘要
Large Language Model (LLM)-based web agents are increasingly evolving from single-agent systems (SAS) to multi-agent systems (MAS). While MAS can lead to improved task performance by decomposing complex tasks across specialized sub-agents, such role decomposition introduces new structural attack surfaces that are absent in SAS. This expanded attack surface remains poorly understood and inadequately categorized. To address this, we propose a taxonomy to categorize attack vectors specific to web-based MAS, accounting for vulnerabilities introduced or amplified by the involvement of multiple agents. We further present a test-bed WebMASLab to analyze web agent security against a fully external, web-only adversary. To isolate the effect of architecture, we keep the user task, tool surface, and browser substrate fixed, and compare single- and multi-agent setups. We evaluate three adversarial scenarios, across three conditions (baseline, prompt-hardened, and reasoning-enabled), including a novel MAS-specific Telephone Loop attack that exploits cross-agent delegation to create cyclical task loops. The attack is inert against SAS but compromises MAS when powered by three of the four frontier models evaluated (Claude Sonnet 4.5, GPT-5.2, GPT-5.4), averaging 80% across them at baseline. Only the fourth model, Claude Sonnet 4.6, resists the attack with a 92% detection rate. For the rest, the detection is 0% at baseline, reaching 33% with prompt-hardening for one model. We also show that obvious defenses do not generalize; prompt-hardening collapses one model's ASR from 100% to 8% while providing only modest reduction to the others. Our findings demonstrate that the transition from single- to multi-agent web systems changes the security landscape. Role specialization may not only lead to performance optimization but also introduce new architectural risks that require further study and defenses.