自适应攻击攻破了针对 LLM 智能体间接提示注入攻击的防御
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Nirma University(尼尔玛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文评估了针对 LLM 智能体间接提示注入攻击的八种防御机制,发现自适应攻击能绕过所有防御并保持超 50% 成功率,揭示了现有防御的脆弱性,强调了引入自适应攻击评估的必要性。
AI中文摘要:
大型语言模型(LLM)智能体通过使用外部工具与环境交互,在各类应用中展现出卓越性能。然而,集成外部工具引入了安全风险,例如间接提示注入(IPI)攻击。尽管针对 IPI 攻击设计了多种防御机制,但由于缺乏对自适应攻击的充分测试,其鲁棒性仍存疑。本文评估了八种不同的防御机制,并利用自适应攻击成功绕过了所有防御,持续实现了超过 50% 的攻击成功率。这揭示了当前防御机制中存在的严重漏洞。我们的研究强调,在设计防御时需要进行自适应攻击评估,以确保系统的鲁棒性与可靠性。代码可在 https://github.com/uiuc-kang-lab/AdaptiveAttackAgent 获取。
英文摘要:
Large Language Model (LLM) agents exhibit remarkable performance across diverse applications by using external tools to interact with environments. However, integrating external tools introduces security risks, such as indirect prompt injection (IPI) attacks. Despite defenses designed for IPI attacks, their robustness remains questionable due to insufficient testing against adaptive attacks. In this paper, we evaluate eight different defenses and bypass all of them using adaptive attacks, consistently achieving an attack success rate of over 50%. This reveals critical vulnerabilities in current defenses. Our research underscores the need for adaptive attack evaluation when designing defenses to ensure robustness and reliability. The code is available at https://github.com/uiuc-kang-lab/AdaptiveAttackAgent.