arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缓解大语言模型中的私人数据泄露:Whiteout

Mitigating Private Data Leakage in LLMs with Whiteout

Anna Yoo Jeong Ha, Ronik Bhaskar, Haitao Zheng, Ben Y. Zhao

arXiv 2610.02418首次发表:更新:

发表机构

University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM记忆并泄露个人敏感信息的问题,本文提出Whiteout工具,通过精确混淆样本覆盖真实PSI,在多种LLM上验证其有效防泄露且对模型效用和安全影响极小,优于现有方法。

AI 中文摘要

现代大语言模型(LLMs)在庞大且大多未经筛选的数据集上进行训练,这些数据集包括从几乎所有可访问网站抓取的内容以及用户输入。因此,LLMs 常常会记忆并重现个人敏感信息(PSI),例如出生日期、电话号码和家庭住址。这带来了显著的隐私风险,尤其是对于高管、政治家和法官等知名人士。现有的缓解措施主要依赖于机器遗忘。然而,这些方法往往删除的信息超出必要范围,降低模型效用和安全性,并且极易受到攻击。本文提出了 Whiteout,一种实用工具,在个人请求下,通过使用精确且精心设计的混淆样本覆盖其真实 PSI,防止 LLMs regurgitating 其真实 PSI。我们在不同规模和制造商的现代 LLMs 上评估了 Whiteout,包括一个广泛使用的 OpenAI 模型。结果表明,Whiteout 有效防止了目标 PSI 的泄露,对模型效用和安全性影响极小,并且优于现有替代方案。我们还针对广泛的对抗措施测试了 Whiteout,从黑盒攻击(如越狱)到白盒自适应攻击(如重新学习和量化)。最后,我们讨论了 Whiteout 的安全和伦理影响。

英文摘要

Modern large language models (LLMs) are trained on massive, largely unfiltered datasets, including content scraped from nearly every accessible website and user inputs. As a result, LLMs often memorize and reproduce personally sensitive information (PSI) such as birth dates, phone numbers, and home addresses. This leads to significant privacy risks, particularly for high-profile individuals such as executives, politicians, and judges. Existing mitigations largely rely on machine unlearning. However, these methods often remove more information than needed, degrade model utility and safety, and are highly vulnerable to attacks. This paper presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine PSIs, by overwriting them using precise and carefully designed obfuscation samples. We evaluate Whiteout on modern LLMs of varying sizes and makers, including a widely-used OpenAI model. Results show that Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, and outperforms existing alternatives. We also test Whiteout against a wide range of countermeasures, from black-box attacks like jailbreaking to white-box adaptive attacks like relearning and quantization. Finally, we conclude with a discussion on the security and ethical implications of Whiteout.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑