arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32293cs.CLcs.AIcs.LG

TRAP:理解并缓解语言模型中的隐私记忆

TRAP: Understanding and Mitigating Privacy Memorization in Language Models

Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu, Holger R. Roth

AI总结:

针对语言模型微调中的隐私记忆问题,提出基于目标参考优势(TRA)的单侧惩罚方法TRAP,在不预先知晓敏感片段的情况下有效降低记忆,且几乎不损失效用。

AI中文摘要:

在敏感记录上微调语言模型可能使其能够复现这些记录。我们探究这种记忆何时产生,以及如何在不预先知道哪些片段敏感的情况下加以预防。我们的出发点是,大多数记忆评分和攻击共享一个统计核心:模型是否赋予某个词元比某个参考更高的概率。以在同一语料库的互补一半上训练的模型作为参考,得到目标参考优势(TRA),这是一种逐词元信号,可将模型对特定记录的拟合与跨记录学习的内容区分开来,且计算成本低、可微。随后我们研究微调期间驱动记忆的因素:记忆在验证最小值之后持续增长,在小数据集和高学习率下更大,在底层任务更难时更高。早停能消除大部分记忆,但由于它是根据总体验证损失选择的,因此对嵌入在原本可学习文本中的稀有、难以预测的片段帮助最小,而这正是敏感信息往往所在之处。因此,我们引入TRAP,一种对逐词元TRA的单侧惩罚,仅在目标模型领先其参考时起作用。在带有标注个人信息的学生论文和带有患者标识符的临床病例上,TRAP将记忆降低到接近未训练模型的水平,而效用损失很小,相比之下,通用正则化器几乎不起作用,差分隐私则放弃了微调带来的大部分收益。

英文摘要:

Fine-tuning a language model on sensitive records can leave it able to reproduce them. We ask when this memorization arises and how to prevent it without knowing in advance which spans are sensitive. Our starting point is that most memorization scores and attacks share one statistical core: whether the model assigns a token more probability than some reference would. Taking as the reference a model trained on the complementary half of the same corpus gives the Target Reference Advantage (TRA), a per-token signal that separates what a model fit to a particular record from what it learned across records, and is cheap and differentiable. We then study what drives memorization during fine-tuning: it keeps growing well past the validation minimum, is larger on small datasets and at higher learning rates, and higher when the underlying task is harder. Early stopping removes much of it, but because it is chosen by aggregate validation loss it helps least for rare, hard-to-predict spans embedded in otherwise learnable text, which is exactly what sensitive information tends to be. We therefore introduce TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference. On student essays with annotated personal information and clinical cases with patient identifiers, TRAP brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.

↑