发表机构
The State Key Laboratory of Blockchain and Data Security, Zhejiang University; Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security; Shanghai AI Laboratory; East China Normal University(浙江大学区块链与数据安全全国重点实验室; 杭州高新技术开发区(滨江)区块链与数据安全研究所; 上海人工智能实验室; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有大语言模型安全评估与防御方法的静态局限,提出达尔文演化攻击-防御框架,含通过策略发现等扩展能力的达尔文攻击及从新对抗样本中迭代学习的达尔文防护,在攻击成功率和防御召回率上表现出色。
AI 中文摘要
大多数现有的大语言模型安全评估和防御方法采用静态形式:用固定攻击方法评估越狱漏洞,在固定恶意提示数据集上训练防护措施。但现实中的对手不断提升能力、扩大攻击空间。为此提出达尔文,一个演化攻击-防御框架,将越狱视为开放式演化过程,通过演化的攻击-防御循环持续更新防护措施。达尔文攻击是通过策略发现、变异、选择和反馈驱动组合来扩展能力的演化对手。在攻击执行时,它根据目标大语言模型和防护措施的反馈自适应选择和组合演化策略。在防御方面,引入达尔文防护,一种在线对抗训练范式,从达尔文攻击生成的新对抗样本中迭代学习。达尔文防护在12个安全基准测试中平均不安全召回率达91.6%,优于如御风-X防护和内莫tron防护等强大防护措施,在标准良性数据集上通过率近100%。
英文摘要
Most existing LLM safety evaluation and defense methods are static: jailbreak vulnerabilities are assessed with fixed attacks, and guardrails are trained on fixed malicious-prompt datasets. In practice, adversaries continually evolve and expand the attack space. We propose DARWIN, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop. DARWIN-Attack evolves its capabilities through strategy discovery, mutation, selection, and feedback-guided composition. It mines strategies from broad external sources, generates variants through self-reflection and genetic evolution, filters them by performance against aligned LLMs, and adaptively selects and composes strategies during attacks. Through continuous evolution, DARWIN-Attack achieves state-of-the-art attack success rates against frontier LLMs and guardrails, reaching nearly 100\% on DeepSeek-V4-Pro, over 90\% on GPT-5.5, and nearly 100\% on YuFeng-XGuard, outperforming recent evolving frameworks such as LSA and MAGIC. Its continued evolution exposes new vulnerabilities, motivating timely defense updates. Accordingly, DARWIN-Guard performs online adversarial training on emerging samples generated by DARWIN-Attack and jointly learns from malicious and benign disguised queries to recognize underlying intent rather than superficial attack patterns. DARWIN-Guard achieves 95.0\% average unsafe recall across multiple safety benchmarks, outperforming recent advanced guardrails such as YuFeng and Nemotron, while maintaining nearly 100\% average benign pass rate on standard benign benchmarks and the best performance on over-refusal benchmarks. Our code and model are available at https://github.com/ZJU-LLM-Safety/DARWIN.