arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoDeL:针对基于LLM的智能体中间接提示注入的协同进化防御

CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

Xiao Yang, Yangchen Ou, Yuhan Gao, Le Wang, Zonghao Ying, Aishan Liu

arXiv 2609.34463首次发表:更新:

AI 中文总结

针对LLM智能体的间接提示注入威胁,提出协同进化防御CoDeL,通过动态重塑攻击分布并联合优化安全与任务完成,显著降低攻击成功率。

AI 中文摘要

基于大语言模型(LLM)的智能体日益依赖外部工具和内容,这使其面临间接提示注入(IPI)的威胁。该威胁催生了多种防御手段,其中基于训练的防御通常被认为最为可靠。然而,现有的基于训练的防御通常在静态的显式注入分布上进行优化,它们学习的是表面形式的线索,而非区分服务用户与服从注入目标之间的边界,因此当恶意意图被融入看似合理的工作流程并延迟数轮执行时,这些防御便会失效。我们提出CoDeL,一种在训练过程中重塑攻击分布并以此强化智能体的防御方法。防御者每轮通过基于LoRA的GDPO,在解耦的奖励(涵盖安全性、任务进展和格式合规性)下进行更新,从而将拒绝注入与完成用户任务共同定义为适应度。为了持续提供值得学习的失败案例,一个协同进化的探测者搜索注入轮次、攻击方法和载荷,寻找仍能穿透当前防御者的注入,并同时以攻击成功率和攻击延迟为指导,优先挖掘防御者过晚察觉的漏洞。每次防御者更新都会使部分攻击种群失效,迫使下一轮探索新的前沿,从而将防御者自身的失败转化为动态课程。在三个IPI基准、九个基线和两个基础模型上的大量实验表明,CoDeL将攻击成功率(ASR)降低了88.5%,并大幅优于其他基线(+38.0%)。代码已公开。

英文摘要

Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑