arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoFlint:多轮大语言模型漏洞的进化图谱

EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

Feitong Qiao, Liren Peng, Shiming Ren, Aishwarya Jadhav, Arghavan Bahadorinejad, Marinette Chen, Muhan Zhang, Abdulaziz Suria, Gennevi Lu, Anish Das Sarma

arXiv 2609.00487首次发表:更新:

发表机构

Reinforce Labs(Reinforce Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出EvoFlint,将进化质量多样性搜索应用于多轮大语言模型红队测试,在HarmBench测试集上对多款模型取得高攻击成功率,生成的结构化攻击图谱可揭示模型安全训练的覆盖情况。

AI 中文摘要

拒绝有害单轮提示的前沿语言模型,当相同意图经多轮逐步达成时往往会顺从,这使得多轮攻击成为大语言模型最不被理解的失效模式之一。大多数自动红队测试方法将其视为生成问题:生成能突破模型的攻击。我们认为将其更好地框架化为搜索问题:发现、组织并迭代优化多样化的攻击策略存档,生成目标模型失效的结构化图谱,而非一次性成功的列表。我们引入EvoFlint,将进化质量多样性搜索应用于多轮红队测试。攻击策略是分阶段的对话计划,而非原始提示,通过大语言模型驱动的变异与交叉进化。基于攻击成功率和峰值严重度的帕累托适应度保留了近失攻击的选择信号。风险索引存档对每个单元内的策略描述嵌入运行带局部竞争的新颖性搜索,在不预设风格分类的情况下维持多样性。代级记忆积累群体中的目标模型洞察并反馈至策略生成。在HarmBench测试集上,EvoFlint在Claude Sonnet 4.6上达到35.8%的攻击成功率,在GPT-5.4上为59.7%,在Qwen3-32B上为94.3%,作为基线参考的旧版GPT-4o则达到98.7%。按风险类别组织的生成存档,为每个目标模型揭示其安全训练已覆盖和未覆盖的伤害类别。

英文摘要

Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation problem: produce attacks that break the model. We argue it is better framed as a search problem: discover, organize, and iteratively refine a diverse archive of attack strategies, producing a structured map of how a target model fails rather than a list of one-off successes. We introduce EvoFlint, which applies evolutionary quality-diversity search to multi-turn red-teaming. Attack strategies are phased conversation plans, not raw prompts, and are evolved through LLM-driven mutation and crossover. A Pareto fitness over attack success rate and peak severity preserves selection signal from near-miss attacks. A risk-indexed archive runs novelty search with local competition over strategy description embeddings inside each cell, maintaining diversity without committing to a predefined style taxonomy. A generation-level memory accumulates target-model insights across the population and feeds them back into strategy generation. On the HarmBench-test split, EvoFlint reaches attack success rates of 35.8% on Claude Sonnet 4.6, 59.7% on GPT-5.4, and 94.3% on Qwen3-32B, alongside 98.7% on the older GPT-4o included as a baseline reference. The resulting archive, organized by risk category, exposes for each target which categories of harm its safety training has and has not covered.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑