边界阻断:审计长时程智能体对抗分阶段提示注入
Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection
浏览论文内容
中文总结 AI 辅助
针对长时程智能体的分阶段提示注入攻击,提出边界行动审计方法,通过路径对齐归因(PAA)在关键行动前预测阻断,在基准上实现更高召回率和更低误阻断率。
中文摘要 AI 辅助
长时程智能体会消费外部内容、调用工具并修改持久状态。间接提示注入可以利用任务特定上下文,在因果关联的阶段间传播,并在工作流继续时改变一个关键行动;我们将此称为分阶段提示注入。我们构建了一个自动化的、反馈引导的攻击生成流水线,并将其应用于Claude Code和Codex的原生运行时。确认的攻击覆盖八个工作流场景、七个攻击目标和六个注入面,表明生产级智能体在长时程中易受上下文感知的多步注入攻击。阻止此类攻击需要在每个关键行动之前做出决策:输入筛选和完成运行评估无法定位干预点,而现有的行动前方法使用不兼容的单位和标签。因此,我们提出了边界行动审计:给定初始上下文、轨迹前缀以及一个完全指定的待处理消息或工具调用,审计器在其效果发生之前预测“通过”或“阻断”。将受攻击与良性执行配对,我们构建了一个包含479对、3112个单元的基准。我们进一步提出了路径对齐归因(PAA),一种无需训练的审计器,它将待处理行动分解为操作元素,并追踪每个值的来源及每个决策的引导因素。PAA仅在模型将某个元素上无根据的实质性影响归因于攻击者可访问的来源时进行阻断,该来源要么提供无条件的引导,要么与可见证据冲突。在完整基准的fail-open评分下,使用Claude Sonnet 5,PAA在6-8%的误阻断率(FBR)下达到86%的阻断召回率,而ARGUS在16-33%的FBR下达到44-47%的召回率。在同一后端下,对于所有三个审计器原生支持的工具调用,PAA的召回率高于VIGIL和ARGUS,且误阻断率更低;所有配对的95%置信区间均排除零。
英文摘要
Long-horizon agents consume external content, invoke tools, and modify persistent state. Indirect prompt injection can exploit task-specific context, propagate across causally connected stages, and alter a consequential action while the workflow continues; we term this staged prompt injection. We build an automated, feedback-guided attack generation pipeline and apply it to Claude Code and Codex in their native runtimes. The confirmed attacks span eight workflow scenarios, seven attack goals, and six injection surfaces, showing that production agents are vulnerable to context-aware, multi-step injection over long horizons. Stopping such attacks requires a decision before each consequential action: input screening and completed-run evaluation cannot locate the intervention point, and existing pre-action methods use incompatible units and labels. We therefore formulate boundary action auditing: given initial context, a trajectory prefix, and a fully specified pending message or tool call, an auditor predicts Pass or Block before its effect occurs. Pairing attacked and benign executions yields a 479-pair, 3,112-unit benchmark. We further propose Path-Aligned Attribution (PAA), a training-free auditor that decomposes pending actions into operative elements and traces what supplied each value and guided each decision. PAA blocks only when the model attributes an unwarranted, material effect on an element to an attacker-reachable source that either provides unqualified steering or conflicts with visible evidence. Under full-benchmark fail-open scoring with Claude Sonnet 5, PAA reaches 86% Block recall at a 6-8% false-block rate (FBR), whereas ARGUS reaches 44-47% recall at 16-33% FBR. Under the same backend, on the tool calls that all three auditors natively support, PAA has higher recall and lower FBR than VIGIL and ARGUS; all paired 95% confidence intervals exclude zero.
发表机构
- Beijing Jiaotong University(北京交通大学)
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
- INRIA Rennes-Bretagne-Atlantique(法国国家信息与自动化研究所雷恩-布列塔尼-大西洋分所)
- Xi’an Jiaotong University(西安交通大学)
机构由 AI 辅助整理,请以论文原文为准。