从定向遗忘模型中提取被遗忘的提示
Extracting Forgotten Prompts from Targeted Unlearned Models
浏览论文内容
中文总结 AI 辅助
该研究发现定向遗忘模型存在新漏洞,提出TAS攻击方法,可从模型中提取被遗忘的提示,在多组实验中实现高准确率且大幅减少查询量。
中文摘要 AI 辅助
近期的遗忘方法(如NPO、DPO、LUNAR)利用拒绝对齐来抑制被遗忘的数据。然而,已有研究表明拒绝响应可能会留下遗忘的痕迹,近期的攻击已能成功恢复部分被遗忘的知识。在本文中,我们发现了一种新的漏洞:现有攻击通常假设攻击者已知被遗忘的提示,其重点是恢复这些提示的答案;而我们证明,可通过保留的数据和对模型的黑盒访问来提取被遗忘的提示本身。我们的攻击方法是定向主动搜索(Targeted Active Search, TAS),它首先通过构建规范模板和实体池,并在有限的查询预算下使用最具信息量的模板-实体对选择性查询模型,来识别被遗忘的实体;一旦识别出实体,TAS就用这些实体实例化提示模板,以探测被遗忘的模型并重建被遗忘的提示。在三种遗忘方法、三个数据集和三个大语言模型(LLM)上开展的实验表明,TAS以100%的准确率恢复了被遗忘的实体,重建了高达95%的被遗忘提示,且查询量比朴素探测少了多达99.7%。
英文摘要
Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data. However, it has been shown that refusal responses might leave traces of unlearning, and recent attacks have been able to successfully recover some of the unlearned knowledge. In this paper, we uncover a new vulnerability. Existing attacks typically assume that the forgotten prompts are already known to the adversary and focus on recovering their answers. However, we show that the forgotten prompts themselves can be extracted by using the retained data and black-box access to the model. Our attack, Targeted Active Search (TAS), first identifies the forgotten entities by constructing canonical templates and entity pool, and selectively querying the model using the most informative template-entity pair under a limited query budget. Once the entities are identified, TAS instantiates prompt templates with those entities to probe the unlearned model and reconstruct the forgotten prompts. Experiments across three unlearning methods with three datasets and three LLMs shows that TAS recovers the forgotten entity with $100\%$ accuracy and reconstructs up to $95\%$ of forgotten prompts, all while using up to $99.7\%$ fewer queries than naive probing.
发表机构
- University of Warwick(华威大学)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。