arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种针对LLM越狱攻击的自进化多智能体框架防御方法

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

Tongyan Hu, Bryan Hooi

arXiv 2608.26008首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM越狱攻击提出一种基于规则记忆的自进化多智能体防御框架,无需参数更新,可降低攻击成功率且保留良性效用,适用于两类LLM模型。

AI 中文摘要

大型语言模型(LLM)仍然易受越狱攻击的影响,这类攻击利用角色扮演、混淆、代码转换和多步间接攻击等技术来诱导模型输出有害内容。随着越狱策略不断涌现,防御措施也在持续发展,形成了一场“猫鼠游戏”,但大多数防御措施是静态的:它们的安全行为在部署时就已固定,无法积累防御经验或适应未见过的策略。我们提出一种基于持久跨交互规则记忆的测试时自进化防御方法:当某次攻击成功时,该框架会将此次失败抽象为一种方法级规则,规则捕获的是攻击包装的结构而非有害主题,并将该规则用于后续输入。由于规则是方法级的,一条诱导出的规则可在整个攻击家族中泛化,且标签空间会随新包装的出现而扩展。该机制完全通过外部记忆和提示实现,无需参数更新,适用于开放权重模型和黑盒API模型。我们将其实现为四个协同模块,但核心贡献是基于记忆的适应机制,而非模块分解。在四个黑盒越狱家族及多个模型上,我们的方法大幅降低了攻击成功率,同时保留了良性效用,在自适应复合包装攻击下仍保持鲁棒性,且不会随记忆增长而增加过度弃权(不执行)的情况。

英文摘要

Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse game, yet most remain static: their safety behavior is fixed at deployment, so they cannot accumulate defensive experience or adapt to unseen strategies. We propose a self-evolving test-time defense built around a persistent, cross-interaction rule memory: when an attack succeeds, the framework abstracts that failure into a method-level rule capturing the structural attack wrapper rather than the harmful topic, and reuses it against future inputs. Because rules are method-level, one induced rule generalizes across an entire attack family, and the label space expands as novel wrappers appear. The mechanism operates entirely through external memory and prompting, with no parameter updates, and applies to both open-weight and black-box API models. We realize it as four cooperating modules, but the contribution is the memory-based adaptation mechanism, not the module decomposition. Across four black-box jailbreak families and multiple models, our method substantially reduces attack success rates while preserving benign utility, remains robust under an adaptive composite-wrapper attack, and does not increase over-refusal as the memory grows.

Comments8 pages (main), with appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑