用户会知道吗?针对使用工具的大语言模型智能体的隐蔽间接提示注入攻击
Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
浏览论文内容
中文总结 AI 辅助
该研究针对工具型LLM智能体的间接提示注入攻击,将攻击成功率分解为隐蔽与公开成功率,提出ICoA诱导隐蔽攻击,在AgentDojo的四个模型上提升了隐蔽成功率3.79-12.01个百分点。
中文摘要 AI 辅助
随着大语言模型(LLM)智能体通过工具执行现实世界操作,间接提示注入(IPI)已成为一项严重威胁。标准评估指标攻击成功率(ASR)仅统计注入是否成功,却忽略了用户在智能体最终响应中注意到的内容。我们对成功的注入轨迹进行分析后发现,存在两种截然不同的结果:智能体执行注入操作,同时返回看似正常的响应;或在最终响应中报告注入的操作,让用户有机会察觉。我们将这两种结果分别称为隐蔽成功和公开成功。从用户视角出发,我们将ASR分解为隐蔽成功率(CSR,统计在最终响应中未留下痕迹的成功注入)和公开成功率(OSR,统计用户可检测到的成功注入)。为理解该差距的驱动因素,我们分析了成功轨迹,发现注入后的智能体行为是区分隐蔽与公开成功的关键:隐蔽轨迹在结束前将控制权交还给用户任务,而公开轨迹则在攻击本身处结束。这种划分遵循ReAct格式,即最终响应会总结最近的操作。基于这一观察,我们提出ICoA(Induced Covert Attack,诱导隐蔽攻击),这是一种旨在通过在执行注入后引导智能体回归用户任务来诱导隐蔽结果的IPI攻击。在AgentDojo上的四个目标模型中,ICoA实现了最高的CSR,相比最强基线提升了3.79至12.01个百分点。
英文摘要
As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agent's final response. Looking at successful injection traces, we find two distinct outcomes: the agent executes the injection while returning an otherwise normal response, or reports the injected action in its final response, giving the user a chance to notice. We call these covert and overt successes. From the user's perspective, we decompose ASR into the Covert Success Rate (CSR), counting successes leaving no trace in the final response, and the Overt Success Rate (OSR), counting successes the user can detect. To understand what drives the gap, we analyze successful trajectories and find that the agent's behavior after the injection separates covert from overt: covert traces hand control back to the user task before ending, while overt traces end at the attack itself. This split follows from the ReAct format, where the final response summarizes the most recent action. Building on this observation, we propose ICoA (Induced Covert Attack), an IPI attack designed to induce covert outcomes by steering the agent back to the user task after executing the injection. Across four target models on AgentDojo, ICoA achieves the highest CSR, with gains of 3.79-12.01 percentage points over the strongest baseline.
发表机构
- Dongguk University-Seoul(东国大学首尔校区)
机构由 AI 辅助整理,请以论文原文为准。