AI 中文总结
该研究发现使用工具的智能体存在框架缺口,间接提示注入攻击可通过重新表述泄露内容突破表层防御,需通过目标约束或能力隔离等手段提升鲁棒性。
AI 中文摘要
持有秘密的使用工具大型语言模型(LLM)智能体在读取攻击者控制的网页内容时,可能遭遇间接提示注入攻击:该内容会迫使智能体泄露秘密。在安全的合成实验室环境(包含金丝雀秘密、模拟工具、匹配的干净与中毒指标)中,我们报告了框架缺口:针对6种模型,10种显性注入类别均被拒绝(GPT-4o为0%),但将相同泄露重新表述为强制性完整性签名、配置字段或类似“可信”主机后,GPT-4o的成功率从0%升至100%。该攻击成本呈三级结构:改写已知机制的措辞极为简单(3种措辞下成功率达96%),在已知有效模板内交换字段也较简单(最高达60%),而围绕新机制生成全新页面则难度较高(0/130)——可复用的资产是模板而非机制。消融实验显示,该机制属于指令与数据混淆,而非对齐防御失效:移除保密政策后,基础攻击成功率仍为0%,重新表述攻击的成功率仅从31.9%升至38.1%。能弥合该缺口的是与载荷无关的检查:当目标被限制时的目标允许列表(成功率0%),以及将规划器与读取器分离的能力隔离方案(成功率0%)。宽泛的“任何形式”政策条款也能在执行模型处弥合缺口(成功率0%),但存在脆弱性(移除通用条款后成功率回升至48.8%)。已发布的微调防御方案SecAlign(CCS 2025)无法在使用工具的智能体上弥合缺口(成功率32.5%,经阳性对照验证),通道分离方案也无法做到(成功率38.8%);输出归一化防护在未见过的编码(如ROT13)面前失效(成功率100%)。鲁棒性源于对目标的约束或能力隔离,而非执行模型对攻击的识别。
英文摘要
A tool-using LLM agent that reads attacker-controlled web content while holding a secret faces indirect prompt injection: the content may make it exfiltrate the secret. In a safe synthetic lab (canary secret, mock tools, matched clean-vs-poisoned metric) we report the framing gap: across six models, ten overt injection classes are refused (gpt-4o 0%), but reframing the identical leak as a mandatory integrity signature, config field, or look-alike "trusted" host drives gpt-4o 0% to 100%. The attack is cheap, and its cost is three-level: paraphrasing a known mechanism is trivial (96% at 3 wordings), swapping the field inside a known-effective template is also cheap (up to 60%), while authoring a fresh page around a new mechanism is hard (0/130) -- the reusable asset is the template, not the mechanism. An ablation shows the mechanism is instruction/data confusion, not defeated alignment: removing the confidentiality policy leaves base attacks at 0% and moves reframing only 31.9% to 38.1%. What closes the gap is payload-blind checks: a destination allow-list (0%, when destinations are closed) and a capability-isolating planner/reader split (0%). A broad "in any form" policy clause also closes it at the acting model (to 0%) but is brittle (dropping the catch-all reopens it to 48.8%). A published fine-tuning defense (SecAlign, CCS 2025) does not close it on a tool agent (32.5%, positive-control-validated), nor does channel separation (38.8%); an output-normalizing guard loses to a held-out encoding (ROT13, 100%). Robustness comes from constraining the destination or isolating the capability, not from the acting model recognizing the attack.
Comments20 pages, 3 figures, 7 tables