arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00392cs.CR

从A2A攻击到包络层防御:LLM智能体的红队评估与三层同构攻防模型

From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model

Yuelin Han

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体的安全评估不足,提出A2A-TIBA攻击原则与GDA测量方法,并构建三层同构攻防模型ELA-ITL,实验验证了攻击有效性与评估有效性。

中文摘要 AI 辅助

ACP和A2A等智能体交互协议已将基于LLM的智能体推向多智能体协作,引入了新的安全威胁。通过A2A由远程对等方发送的任务被视为合法请求,为间接提示注入提供了自然通道。现有的智能体安全评估大多依赖单一指标——攻击成功率(ASR),且无法区分攻击失败是因为LLM识别了恶意内容,还是因为智能体层的机制阻止了执行。为解决这一问题,我们提出A2A-TIBA,一种结合间接提示注入与绕过规避的攻击原则。通过植入-命令-窃取步骤,它诱导目标智能体部署回调交互程序,之后攻击者发出绕过智能体的命令。为更精细地评估防御,我们设计了GDA测量,一种红队测试平台方法,通过LLM网关进行原始上下文捕获、双重数据保留和基于智能体的自主评判。我们提出四种攻击结果,A/B/C/D类,将ASR扩展为语义拒绝率、语义突破率、拦截率和渗透率。实验揭示了包络层——恶意内容进入智能体的通道——作为一个新的防御维度。我们据此提出ELA-ITL,一种三层同构攻防模型,将防御分为包络封装、LLM识别和智能体拦截,将攻击分为植入通道、提示优化和执行机制。在15种智能体前端×LLM后端组合上的测试以及构建1000案例数据集,验证了A2A-TIBA的攻击有效性、GDA测量的评估有效性,并确认在包络封装(如A2A、工具和内存通道)中添加恶意提示标签显著提高了LLM对恶意内容的识别。

英文摘要

Agent interaction protocols such as ACP and A2A have moved LLM-based agents toward multi-agent collaboration, introducing new security threats. A task sent by a remote peer over A2A is treated as a legitimate request, providing a natural channel for indirect prompt injection. Existing agent security evaluations mostly rely on a single metric, the attack success rate (ASR), and cannot distinguish whether an attack failed because the LLM recognized the malicious content or because a mechanism at the agent layer blocked execution. To address this, we propose A2A-TIBA, an attack principle combining indirect prompt injection with bypass circumvention. Through implant-command-exfiltration steps, it induces the target agent to deploy a callback interaction program, after which the attacker issues commands bypassing the agent. To evaluate defenses finer, we design GDA Measurement, a red-team testbed method using raw context capture via an LLM gateway, dual data preservation, and agent-based autonomous judging. We propose four attack outcomes, Class A/B/C/D, extending ASR into semantic refusal rate, semantic breach rate, interception rate, and penetration rate. Experiments reveal the envelope layer -- the channel through which malicious content enters an agent -- as a new defense dimension. We accordingly propose ELA-ITL, a three-layer isomorphic attack-defense model, dividing defense into envelope packaging, LLM recognition, and agent interception, and attack into implant channel, prompt optimization, and execution mechanism. Testing on 15 agent front-end x LLM back-end combinations and building a 1,000-case dataset verifies the attack effectiveness of A2A-TIBA, the evaluation validity of GDA Measurement, and confirms that adding malicious prompt labels to envelope packaging such as A2A, tool, and memory channels significantly improves LLM recognition of malicious content.

↑