AI 中文总结
本文提出一种主动防御框架,通过协作多智能体架构结合干扰、误导与自适应策略,在EMRA数据集上使LLMs对抗多轮攻击的ASR平均降低69%,同时提升DR与攻击者token消耗。
AI 中文摘要
随着大型语言模型(LLMs)越来越多地融入复杂应用,其对抗性攻击的脆弱性引发了重大担忧。然而,现有防御本质上仍是被动式的,这一局限使其难以应对复杂威胁,因为攻击者会在多轮交互中持续调整策略。本文提出一种主动防御框架,用于保护LLMs抵御演进式多轮对抗性攻击,该框架在连续交互轮次中结合了干扰、误导与自适应策略。具体而言,它采用协作多智能体架构,其中专门的智能体执行互补的防御策略,包括受控响应 pacing(节奏)以增加攻击成本、策略性模糊输出以误导攻击者采取无效策略,以及对交互日志的取证分析以识别攻击模式并完善防御。这些智能体由自适应机制协调,该机制会根据升级的威胁动态调整防御策略。为便于全面评估,本文提供了EMRA数据集,旨在模拟多轮攻击中的演进策略,包含8种攻击类型的5200个对抗样本。在EMRA数据集上针对多个LLM主干的实验结果显示,与评估的最先进基线相比,该框架平均降低了69%的攻击成功率(ASR);除了抑制有害输出外,它还维持了欺骗性交互,平均欺骗率(DR)是最强基线的6倍以上,且相对于评估基线平均增加了198.83%的攻击者 token 消耗。代码和数据集可在指定URL获取。
英文摘要
As LLMs become increasingly integrated into complex applications, their vulnerability to adversarial attacks has raised significant concerns. However, existing defenses remain reactive in nature. This limitation makes it difficult for them to counter sophisticated threats, as adversaries continuously adjust their strategies across multi-turn interactions. In this paper, we present a proactive defense framework for securing LLMs against evolving multi-turn adversarial attacks that combines disruption, misdirection, and adaptation across successive interaction turns. In particular, it employs a cooperative multi-agent architecture in which specialized agents execute complementary defense strategies. These strategies include controlled response pacing to increase attack costs, strategically ambiguous outputs to mislead adversaries into ineffective strategies, and forensic analysis of interaction logs to identify attack patterns and refine defenses. These agents are coordinated by an adaptive mechanism that dynamically adjusts the defense strategy in response to escalating threats. To facilitate comprehensive evaluation, we present the EMRA dataset designed to simulate evolving strategies across multi-turn attacks, including 5,200 adversarial samples across eight attack types. Experimental results on EMRA across multiple LLM backbones show that the proposed framework reduces ASR by 69% on average relative to evaluated state-of-the-art baselines. Beyond suppressing harmful outputs, it sustains deceptive engagement, achieving an average DR more than six times that of the strongest baselines and increasing attacker-token consumption by 198.83% on average relative to evaluated baselines. Code and dataset are available at https://github.com/SiyuanLi00/CoopGuard.