APEX:面向LLM智能体执行边界的主动防护
APEX: Active Protection at Execution Boundaries for LLM Agents
- Tsinghua University(清华大学)
- Imperial College London(帝国理工学院)
- Nanjing University(南京大学)
- The Chinese University of Hong Kong(香港中文大学)
- University College London(伦敦大学学院)
- A*STAR(新加坡科技研究局)
- University of British Columbia(不列颠哥伦比亚大学)
- Shenzhen University(深圳大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
APEX通过在执行边界强制执行基于授权契约的证据门控预防和欺骗暴露,统一防护LLM智能体免受间接提示注入,在多个基准上实现近零攻击成功率。
AI中文摘要:
间接提示注入(IPI)将对抗性指令隐藏在大语言模型(LLM)智能体在运行时读取的内容中。随着智能体组合异构能力单元(包括工具、MCP服务器和技能),注入的载体成倍增加,而旨在识别攻击模式的防御已落后于它们。我们转而将防御从覆盖攻击模式转移到单一稳定点:无论载体是什么,无论注入如何传播,危害只会在执行边界(智能体将内部状态转化为外部动作或释放输出的地方)产生。那里的安全取决于两个条件,两者都由受信任的任务而非运行时决定:提议的效果是否被授权,以及到达该效果的运行时信息是否得到该任务的认可。我们提出APEX,一种主动防御,在不可信执行之前编译的单一授权契约下,在此边界强制执行这两个条件:证据门控预防仅在契约证明效果合理时才允许该效果,而基于欺骗的暴露使未经认可的使用在效果提交之前自我暴露。因此,防护源于任务允许的内容,而非攻击的构建方式,并且统一适用于所有能力单元,无需特定于攻击的策略或污点跟踪。与13个基线相比,APEX在六个基准中的五个上实现了0%的攻击成功率,在第六个上为0.56%,在所有三种能力单元类型上在自适应攻击下保持0%,并在不同防御者骨干下保持有效。代码可在https URL获取。
英文摘要:
Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind them. We instead shift defense from covering attack patterns to one stable point: whatever the carrier and however the injection propagates, harm materializes only at the \emph{execution boundary}, where the agent turns internal state into an external action or released output. Safety there turns on two conditions, both settled by the trusted task rather than by the run: whether the proposed effect is authorized, and whether the runtime information reaching it is endorsed by that task. We present APEX, an active defense that enforces both at this boundary from a single authorization contract compiled before untrusted execution: \emph{evidence-gated prevention} admits an effect only when the contract justifies it, while \emph{deception-based exposure} makes unendorsed use reveal itself before the effect commits. Protection therefore follows from what the task permits rather than from how an attack is built, and applies uniformly across capability units without attack-specific policies or taint tracking. Against 13 baselines, APEX attains 0\% attack success on five of six benchmarks and 0.56\% on the sixth, holds 0\% under adaptive attacks on all three capability-unit types, and remains effective across defender backbones. Code is available at https://github.com/ZhengXR930/APEX_official/tree/official.