AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
AutoAdv:多轮对抗性提示生成用于大型语言模型的多轮 Jailbreaking 攻击
专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(title);LLM(abstract)
AI总结 AutoAdv 提出了一种自动化多轮对抗性提示生成方法,通过策略性重写和优化配置,实现对大型语言模型的安全机制的高效攻击,揭示了其在有害内容生成上的高成功率。
Comments We encountered issues with the paper being hosted under my personal account, so we republished it under a different account associated with a university email, which makes updates and management easier. As a result, this version is a duplicate of arXiv:2511.02376