发表机构
University of Illinois Urbana-Champaign; University of Ottawa; Tongji University(伊利诺伊大学厄巴纳-香槟分校; 渥太华大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对大型语言模型的私有信息遗忘问题,构建含虚假私有信息的合成数据集,提出白盒审计框架,发现五种现有遗忘方法中,逆贪心解码可恢复所谓已遗忘的信息,表明当前遗忘方法常无法完全消除敏感信息,需更可靠方法保障LLMs隐私。
AI 中文摘要
大型语言模型(LLMs)会记忆敏感信息,引发严重的隐私问题。机器遗忘提供了一种移除这类信息的潜在解决方案,但现有方法是真正擦除了信息,还是仅将其隐藏在模型内部,目前仍不明确。一个关键挑战是在统一评估框架下量化敏感数据的留存程度。为解决这一问题,我们构建了包含虚假私有信息的合成数据集,并提出一种白盒审计框架,以系统评估声称已被遗忘的信息是否被真正移除。利用该框架,我们评估了五种现有的遗忘方法,发现一种简单的“逆贪心”解码(每一步选择概率最低的token)能够恢复所谓已被遗忘的私有信息。我们的结果表明,当前的遗忘方法往往无法完全消除敏感信息,凸显出需要更可靠的方法来确保部署的LLMs的隐私安全。
英文摘要
Large language models (LLMs) can memorize sensitive information, raising serious privacy concerns. Machine unlearning offers a potential solution to remove such information, but it remains unclear whether existing methods truly erase it or merely hide it within the model. A key challenge is quantifying the persistence of sensitive data under a unified evaluation framework. To address this, we construct a synthetic dataset containing fake private information and propose a white-box auditing framework to systematically assess whether claimed-forgotten information is genuinely removed. Using this framework, we evaluate five existing unlearning methods and find that a simple "inverse greedy" decoding -- selecting the least likely token at each step -- can recover supposedly forgotten private information. Our results reveal that current unlearning approaches often fail to fully eliminate sensitive information, highlighting the need for more reliable methods to ensure privacy in deployed LLMs.