arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16098cs.CRcs.AIcs.CL

面向工具集成大语言模型智能体的通用对抗攻击防御方法

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

Xiaoyan Li, Yunli Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对工具集成LLM智能体面临的四类对抗攻击,提出两种通用工具防御(攻击者工具过滤和正常工具召回)及提示防御,在多种开源和专有模型上将攻击成功率降至0%,同时保持或提升任务成功率。

中文摘要 AI 辅助

大语言模型(LLM)智能体在多个领域展现出令人印象深刻的能力,尤其是在与外部工具集成以完成多步任务时。然而,它们越来越容易受到对抗性攻击,包括直接提示注入、间接提示注入、记忆投毒和后门攻击,这些攻击利用了模型对提示注入和工具操作的开放性。在这项工作中,我们在一个统一框架内探索了针对这四类攻击的实用且可泛化的防御策略。我们引入了两种通用的基于工具的防御方法:攻击者工具过滤,它使用异常检测(例如孤立森林)来识别并移除可疑工具;以及正常工具召回,一种白盒方法,在规划之前恢复智能体的原始工具集。此外,我们纳入了基于提示的防御:思维链提示和自反思技术以增强推理,以及任务改写以缓解攻击。在四个开源LLM(Gemma2-9B、Qwen2-7B、LLaMA3-8B和LLaMA3.1-8B)和三个专有LLM(GPT-3.5、GPT-4和GPT-5)上的实验结果表明,我们的方法显著降低了攻击成功率(ASR),在许多设置中实现了0%的ASR,同时保持甚至提高了原始任务成功率。这些发现凸显了简单、模块化、多层防御在增强工具集成LLM智能体的安全性和鲁棒性方面的前景。代码可在以下网址获取:https://this https URL。

英文摘要

Large Language Model (LLM) agents have demonstrated impressive capabilities across a variety of domains, particularly when integrated with external tools for multi-step task completion. However, they are increasingly vulnerable to adversarial attacks, including direct prompt injection, indirect prompt injection, memory poisoning, and backdoor attacks, which exploit the model's openness to prompt injection and tool manipulation. In this work, we explore practical and generalizable defense strategies within a unified framework across these four attack types. We introduce two universal tool-based defenses: Attacker Tool Filtering, which uses anomaly detection (e.g., Isolation Forest) to identify and remove suspicious tools, and Normal Tool Recalling, a white-box method that restores the agent's original toolset prior to planning. Additionally, we incorporate prompt-based defenses: Chain-of-Thought prompting and self-reflection techniques to enhance reasoning and task paraphrasing to mitigate attacks. Experimental results across both four open-source LLMs (Gemma2-9B, Qwen2-7B, LLaMA3-8B, and LLaMA3.1-8B) and three proprietary LLMs (GPT-3.5, GPT-4, and GPT-5) show that our methods significantly reduce the Attack Success Rates (ASR), achieving 0% ASR in many settings, while preserving or even improving the original task success rate. These findings highlight the promise of simple, modular, multi-layered defenses for strengthening the security and robustness of tool-integrated LLM agents. The code is available at https://github.com/Xiaoyan-Lisa/Defenses-for-Tool-Integrated-LLM-Agents-Against-Adversarial-Attacks.

发表机构

  • University of Toronto(多伦多大学)
  • National Research Council Canada(加拿大国家研究委员会)

机构由 AI 辅助整理,请以论文原文为准。

↑