AI 中文总结
该研究针对LLMs,提出结合关联上下文检索的知识编辑白盒攻击,提升攻击有效性且不严重损害模型整体性能。
AI 中文摘要
随着大语言模型(LLMs)被赋予越来越多的自主权,研究可诱导其不安全行为的方法至关重要。我们提出一种受知识编辑领域“定位后编辑”方法启发的新型白盒攻击,选择该方法的依据是观察到采用此类方案编辑后的模型往往会对编辑目标分配异常高的预测概率,这一特性在设计攻击时尤为有利。我们通过纳入从模型中检索到的关联知识来修改编辑框架,从而将约束移除扩展至整个主题类别,而非局限于预定义数据集的提示词。对多种架构开展的实验表明,与竞争方法相比,该攻击的有效性有所提升,且不会对模型的整体性能造成严重损害。
英文摘要
As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction probabilities to the edit target, a property that is particularly advantageous when designing attacks. We modify the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset. Experiments with various archi- tectures demonstrate improved attack effectiveness over competing methods without dealing critical damage to general model performance.