Attributing and Exploiting Safety Vectors through Global Optimization in Large Language Models
通过全局优化在大语言模型中归因和利用安全向量
机构 * Southeast University(东南大学) ; Zhejiang University(浙江大学) ; OPPO Research Institute(OPPO研究院)
AI总结 通过全局优化方法识别大语言模型中的安全向量,揭示安全机制的协作交互,并提出新的白盒劫持方法提升攻击效果。