Jailbreak Attack Initializations as Extractors of Compliance Directions
越狱攻击初始化作为合规方向的提取器
机构 * Department of Computer Science, Technion - Israel Institute of Technology(技术学院计算机科学系) ; Department of Data and Decision Science, Technion - Israel Institute of Technology(技术学院数据与决策科学系) ; School of Electrical and Computer Engineering Engineering, Ben-Gurion University of the Negev(内盖夫本· Gurion大学电气与计算机工程学院)
AI总结 本文发现基于梯度的越狱攻击初始化会收敛到抑制拒绝的单一合规方向,并据此提出CRI框架,通过沿合规方向投影未见提示来提高攻击成功率并降低计算开销。
Comments Accepted to Findings of the Association for Computational Linguistics 2025 (EMNLP 2025)