发表机构
Chongqing University; Zhejiang University; Huazhong University of Science and Technology(重庆大学; 浙江大学; 华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无数据攻击ReGap,利用LLM生成候选和低秩适应恢复私有关联,在多个模型上提升恢复率6-21个百分点,揭示微调可重新唤醒隐私风险。
AI 中文摘要
除了将大型语言模型(LLMs)适应于专门应用之外,最近研究表明,微调可以恢复通过直接查询不再可访问的私有信息。然而,先前的微调恢复攻击需要来自同一分布的真实私有监督,即先前的训练数据集。我们认为,在没有这种不切实际的知识的情况下,这种恢复仍然是可能的。我们表明,LLM生成的候选可以提供足够的监督来恢复先前学习的私有关联。基于此,我们提出了ReGap,一种无数据的攻击,利用任务结构恢复私有关联,通过答案标记似然过滤它们,并通过低秩适应更新目标模型。具体来说,ReGap既不需要目标答案,也不需要辅助的真实私有监督。在六个GPT-2、OPT和Qwen3模型中,ReGap在目标关联恢复上比后训练目标模型提高了6-21个百分点。即使适应身份与所有记忆和评估身份不相交,且生成或选择的监督中不出现确切的目标答案,恢复仍然显著。此外,相同的训练适配器在先前暴露的检查点上将恢复率从42%提高到63%,但在从未遇到目标的匹配检查点上没有产生增益。这种对比表明,仅适应不足以解释观察到的恢复,先前的目标暴露强烈影响适应后的可恢复性。我们的发现强调,常规模型定制可以重新唤醒潜在的隐私风险,值得学术界和工业界紧急关注。
英文摘要
Beyond adapting Large Language Models (LLMs) to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same distribution, i.e., the previous training dataset. We argue that such recovery remains possible without such impractical knowledge. We show that LLM-generated candidates can provide sufficient supervision to recover previously learned private associations. Based on this, we propose ReGap, a data-free attack that recovers private associations using task structure, filters them by answer-token likelihood, and updates the target model via low-rank adaptation. Specifically, ReGap requires neither target answers nor auxiliary genuine private supervision. Across six GPT-2, OPT, and Qwen3 models, ReGap improves target-association recovery by 6-21 percentage points over the post-training target model. Recovery remains substantial even when the adaptation identities are disjoint from all memorized and evaluation identities, with no exact target answers appearing in the generated or selected supervision. Moreover, the same trained adapters increase recovery from 42\% to 63\% on a previously exposed checkpoint, but produce no gain on a matched checkpoint that never encountered the targets. This contrast shows that adaptation alone is insufficient to explain the observed recovery and that prior target exposure strongly affects post-adaptation recoverability. Our findings highlight that routine model customization can reawaken latent privacy risks, warranting urgent attention from the academic and industrial communities.
Comments34 pages