arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21727cs.LGcs.AI

基于良性事实的强化学习会放大模型对已记忆私人数据的泄露

Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data

Renfei Zhang, Niloofar Mireshghallah

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现,基于良性事实的强化学习可放大指令模型对已记忆个人身份信息的泄露,且推理能力与拒绝率得以保留,为攻击者提供了无需接触私人数据即可提取的途径。

中文摘要 AI 辅助

带可验证奖励的强化学习(RLVR)被用于提升模型的推理任务能力,但其对模型会泄露什么内容的副作用尚未得到充分研究。本文表明,基于事实的强化学习会提升指令模型对已记忆的个人身份信息(PII)的提取能力。我们首先确认指令模型已记忆了PII,但这些信息处于潜伏状态,在被询问时很少会浮现出来。随后,我们对完全不包含任何PII的良性事实数据应用强化学习,并重新进行探测:包括针对姓名-邮箱对的定向探测,以及仅要求模型列出其已知地址的无定向自由回忆提示。两种探测下的PII提取率均大幅上升:在DeepSeek-V3.1上,逐字回忆@k从0.155提升至0.370,增幅达2.4倍。该效应随模型规模扩大而增强:在参数规模从80亿到6710亿的三个模型中,最大的模型出现了绝对泄露量最高的情况。与此同时,模型的推理能力和拒绝率得以保留,表明强化学习仅选择性地改变了已记忆信息的可访问性,而非广泛改变模型本身。综上,通过从未接触过私人数据的训练,可使已记忆的私人数据显著更易被提取。这为攻击者提供了一条无需隐私相关训练信号、无需访问私人数据本身,仅需对无害内容进行微调即可获取已记忆数据的途径。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally identifiable information (PII) the instruct model had already memorized. We first confirm that instruct models have already memorized PII but leave them latent, rarely surfacing one when asked. We then apply RL on benign factual data that contains no PII of any kind, and re-probe: a targeted probe over name->email pairs, and an untargeted free-recall prompt that simply asks the model to list the addresses it knows. PII extraction rises sharply under both: on DeepSeek-V3.1, verbatim recall@k increases from 0.155 to 0.370, a 2.4x gain. The effect scales with model size: across three models spanning 8B to 671B parameters, absolute leakage is largest in the biggest model. Meanwhile model's reasoning abilities and refusal rates are retained, indicating that RL selectively changes which memorized information is accessible rather than broadly altering the model. In summary, memorized private data can be made markedly more extractable by training that never touches it. This gives an adversary a route to memorized data that requires no privacy-relevant training signal and no access to the data itself -- only the ability to fine-tune on something innocuous.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑