arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JailbreakSkill:通过可复用且不断演进的技能扩展自动化红队测试

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

Xiaoyu Wen, Jiajia Li, Zhida He, Peng Yu, Chenxu Wang, Han Qi, Ziyuan Zhou, Cheng Jin, Ying Wen, Xingcheng Xu, Shuyue Hu, Tianhang Zheng, Chaochao Lu, Qiaosheng Zhang

arXiv 2608.16465首次发表:更新:

发表机构

Shanghai AI Laboratory; Northwestern Polytechnical University; Fudan University; Shanghai Jiao Tong University; Zhejiang University(上海人工智能实验室; 西北工业大学; 复旦大学; 上海交通大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出以技能为核心的JailbreakSkill框架,将攻击策略打包为可复用技能,通过攻击与学习的闭环演进技能,大幅提升攻击成功率,且技能可泛化至未见过的提示词与目标模型。

AI 中文摘要

自动化红队测试已产生越来越多的攻击策略,但这些策略通常分散在提示词和工作流程中,难以大规模地进行系统整合、复用和改进。我们推出JailbreakSkill,这是一个以技能为核心的框架,通过可复用且持续演进的攻击能力来扩展自动化红队测试。JailbreakSkill将现有攻击策略打包为模块化、可用于智能体的技能,这些技能可在不同任务和目标模型中直接复用和自适应选择。除复用外,它还打通了攻击与学习的闭环:攻击经验被用于诊断、优化、组合和发现新技能,这些新技能会被添加到不断增长的技能库中。这种演进使AdvBench上的宏观平均攻击成功率(ASR)提升了17.5个百分点,HarmBench上提升了13.4个百分点,其中针对GPT-5.4在AdvBench上的攻击成功率提升了48.6个百分点,同时还产生了新的攻击策略,例如将直接请求重新构建为未完成的文档补全任务。部分演进后的技能无需进一步适配即可泛化到未见过的提示词和目标模型。我们的代码可在此https URL获取。

英文摘要

Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale. We introduce \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities. \textsc{JailbreakSkill} packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target models. Beyond reuse, it closes the loop between attacking and learning: attack experience is used to diagnose, refine, combine, and discover new skills, which are added back to an ever-growing skill library. This evolution lifts macro-average ASR by 17.5 percentage points on AdvBench and 13.4 points on HarmBench, including a 48.6-point gain against GPT-5.4 on AdvBench, while yielding novel attack strategies such as reframing a direct request as an unfinished document-completion task. Several evolved skills also generalize to unseen prompts and target models without further adaptation. Our code is available at https://github.com/BattleWen/JailbreakSkill.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑