arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用知识编辑中的关联上下文检索构建针对大语言模型(LLMs)的白盒攻击

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

Roman Maksimov, Vladimir Aletov, Vladimir Solodkin, Dmitry Bylinkin, Daniil Medyakov, Aleksandr Beznosikov

arXiv 2608.17836首次发表:更新:

AI 中文总结

该研究针对LLMs,提出结合关联上下文检索的知识编辑白盒攻击,提升攻击有效性且不严重损害模型整体性能。

AI 中文摘要

随着大语言模型(LLMs)被赋予越来越多的自主权,研究可诱导其不安全行为的方法至关重要。我们提出一种受知识编辑领域“定位后编辑”方法启发的新型白盒攻击,选择该方法的依据是观察到采用此类方案编辑后的模型往往会对编辑目标分配异常高的预测概率,这一特性在设计攻击时尤为有利。我们通过纳入从模型中检索到的关联知识来修改编辑框架,从而将约束移除扩展至整个主题类别,而非局限于预定义数据集的提示词。对多种架构开展的实验表明,与竞争方法相比,该攻击的有效性有所提升,且不会对模型的整体性能造成严重损害。

英文摘要

As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction probabilities to the edit target, a property that is particularly advantageous when designing attacks. We modify the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset. Experiments with various archi- tectures demonstrate improved attack effectiveness over competing methods without dealing critical damage to general model performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑