arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27782cs.CRcs.CLcs.LG

记忆并非提取:严格的差分隐私边界与审计盲区

Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots

Xujun Che, Depeng Xu, Shuhan Yuan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究明确了大型语言模型中反事实记忆与自适应提取的差分隐私边界,发现二者互不控制,存在双向盲区,且在数十亿参数模型上依然存在。

中文摘要 AI 辅助

大型语言模型中的记忆通过大量定义来衡量,这些定义之间的形式关系尚不明确,且差分隐私(DP)被同时当作抵御所有这些定义的代理。我们确定了两种具有实际重要性的定义——反事实记忆和自适应提取的精确DP常数,并表明它们互不控制。在f-DP下,对于无知基线κ,任何具有列表预算m的自适应提取协议成功的概率至多为1-f(κ),且该边界在密集基线集合上是紧的:DP在秘密可先验猜测的程度上,恰好控制提取直至某个阈值。最小熵证明基线与分布无关,因为在纯ε-DP下,对于任何先验,当提取风险水平τ≤1/2时,满足H_∞≥εlog₂e+log₂(m/τ),且在均匀先验下该等式精确成立。在记忆方面,f-DP将任何有界分数的反事实记忆限制在优势函数η(f)内,在纯DP下该函数等于tanh(ε/2);对于k≥2个重复副本,朴素的ε→kε边界tanh(kε/2)无法达到,精确常数是由几何噪声计数得到的闭式阶梯函数。该限制在实际使用的局部分数类中达到,正是在此处两种度量分离:一种机制被记忆但不可提取,另一种完全可提取但对所有基于损失的分数完全不可见。这为基于损失的审计和遗忘验证打开的双向盲区在数十亿参数模型上依然存在:从一个提示中可逐字恢复保留触发内容,而从业者部署的审计却证明其干净。

英文摘要

Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is treated as a proxy against all of them at once. We pin down the exact DP constant for the two that carry the practical weight, counterfactual memorization and adaptive extraction, and show that they do not control each other. Under $f$-DP, every adaptive extraction protocol with list budget $m$ succeeds with probability at most $1-f(κ)$ for the oblivious baseline $κ$, and the bound is tight on a dense set of baselines: DP uniformly controls extraction exactly up to a threshold in how well the secret can be guessed a priori. Min-entropy certifies that baseline distribution-free, since $H_\infty\geε\log_2 e+\log_2(m/τ)$ holds extraction below a risk level $τ\le1/2$ under pure $ε$-DP for every prior, and is exact on uniform priors. On the memorization side, $f$-DP caps the counterfactual memorization of any bounded score at an advantage functional $η(f)$, equal to $\tanh(ε/2)$ under pure DP; for $k\ge2$ duplicated copies the naive $ε\mapsto kε$ bound $\tanh(kε/2)$ is unattainable, the exact constant being a closed-form staircase attained by geometric noisy counting. That cap is attained inside the local score class used in practice, and it is there that the two measures separate: one mechanism is memorized yet unextractable, another fully extractable yet exactly invisible to every loss-based score. The two-sided blind spot this opens for loss-based auditing and unlearning verification survives on billion-parameter models: a reserved-trigger release is recovered verbatim from one prompt while the audits practitioners deploy certify it clean.

↑