Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
针对对齐:提取对齐LLM的安全分类器
机构 * 2 Department of Computer Science Virginia Tech Blacksburg, VA, USA
专题命中 越狱攻击 :alignment(title,abstract);safety(title,abstract);jailbreak(abstract);分类 cs.AI
AI总结 本文提出了一种提取对齐LLM安全分类器的方法,通过构建候选分类器并评估其在对抗性攻击中的有效性,展示了替代分类器在提高攻击成功率方面的优势。
Comments This work has been accepted for publication at the IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). The final version will be available on IEEE Xplore