发表机构
School of Software Engineering, South China University of Technology(华南理工大学软件学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工具增强语言代理的间接提示注入,提出基于验证器的协同进化强化学习框架CoRL,通过攻击者SFT、双边Co-PPO和防御者SFT三阶段训练,将攻击成功率降至0.0%并提升效用至76.3%。
AI 中文摘要
工具增强的语言代理容易受到间接提示注入(IPI)的攻击。与直接提示注入不同,IPI将对抗性指令隐藏在不受信任的工具输出中,能够隐蔽地改变合法任务的执行。针对固定攻击训练的防御可能在攻击者改变其策略、注入位置和载荷时失效。为解决这一问题,我们将自适应IPI建模为非对称、部分可观测的一般和马尔可夫博弈:多轮攻击者根据公开轨迹在到达的工具返回位置自适应调整载荷,而使用工具的防御者必须阻止注入目标并完成用户任务。我们提出CoRL,一个基于验证器的协同进化与修复框架,包含三个阶段:攻击者SFT从成功轨迹初始化多轮攻击;双边Co-PPO使用角色特定奖励和历史对手种群联合训练两个代理;防御者SFT整合验证器接受的教师修复以应对种群发现的失败。在每位防御者的1,514次干净、固定模板和自适应执行中,CoRL将整体攻击成功率(ASR)降低38.5个百分点至0.0%,并将效用提升13.1个百分点至76.3%。分阶段和受控消融实验表明在线Co-PPO和种群挖掘修复具有积极贡献,而外部基准评估表明攻击抵抗具有迁移性。在所评估的攻击下,防御者在安全性和任务效用之间取得平衡,而保留的攻击者为自适应红队评估提供候选。
英文摘要
Language-model agents are vulnerable to indirect prompt injection (IPI) during tool use: adversarial instructions hidden in untrusted tool outputs can covertly redirect legitimate task execution. Existing work often trains and evaluates defenses against fixed attacks that do not adapt to the defender's behavior, so the resulting defenses may struggle against adaptive attacks in real-world settings. We argue that a strong defense against adaptive IPI must adapt during training to a continually evolving attacker. Building on this insight, we propose CoER, a verifier-grounded co-evolution and refinement framework that models interleaved tool calls and adaptive injections within a task as a general-sum Markov game: the defender advances the task through successive tool calls, while the attacker can inject multiple times within the same task and adapt subsequent attacks to the defender's responses and prior execution traces. After initializing the attacker from successful trajectories, bilateral adversarial reinforcement learning (BA-RL) retains historical policies from both roles as opponent populations and mixes current and historical opponents, extending training beyond the latest matchup. Attackers from these populations are then reused to challenge teacher agents, and only demonstrations verified for both safety and task completion are used to fine-tune the co-evolved defender. Across seven domains and three evaluation seeds, CoER reduces adaptive attack success from 41.3% to 0.2% and raises safe task completion from 39.6% to 76.2%; external benchmarks also show improved attack resistance. Further experiments validate the effectiveness of bilateral historical-opponent mixing and population-guided refinement. Attacker analyses show that co-evolution strengthens attack capabilities and that the trained attacker uses execution feedback to adapt subsequent injections.
Comments32 pages, 7 figures; synchronized with the current ICLR manuscript. Project page: https://ilianzby.github.io/CoER/