PAC私有自回归生成:将噪声校准到集成分歧
PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement
浏览论文内容
中文总结 AI 辅助
本研究将PAC隐私从分类扩展到自回归生成,通过构建多个世界并校准噪声至集成分歧,在保持隐私预算的同时显著优于PMixED,保留大部分非私有性能。
中文摘要 AI 辅助
在私有文本上适配的语言模型通常通过API提供服务,因此隐私泄露发生在生成的输出中,而非暴露的权重。私有预测保护了这些发布。诸如PMixED之类的方法在每次发布时都会产生隐私成本,并且在长时间范围内越来越依赖公共模型。PAC隐私则根据可能秘密之间的输出变异性来校准噪声,当预测稳定时添加较少的噪声。据我们所知,PAC私有预测此前尚未从分类扩展到自回归生成。我们从私有语料库构建$m=128$个重叠的世界,每条记录恰好出现在$m/2$个世界中,并在冻结的公共模型上为每个世界训练一个适配器。实际世界即为秘密。在每个token处,公共模型定义候选集,各世界进行投票,其后验加权分歧决定PAC噪声;全体一致时无需校准噪声。我们证明$I(S;Y_{1:T}) \leq I(S;H_T) \leq bT$。我们的贡献是将PAC隐私扩展到自回归生成,处理自适应自生成上下文,并引入耦合解码,在保持隐私核算的同时避免贪婪退化。在WikiText-103上使用GPT-2-small,我们在每token预算为$2^{-32}$时保留了微调收益的74%,而成员推断成功率在$10^6$个token后被限制在51.08%;后验熵估计的泄漏约为收费预算的17%。推断隐私不是内容保护:即使对记忆的金丝雀的成员优势与零无法区分,金丝雀仍以相同速率被发出。与PMixED在相同数据宇宙和测试集上匹配成员推断界限相比,我们在$10^2$到$10^6$个token范围内保留了非私有余量的98%,而PMixED最多保留56%,且无交叉。
英文摘要
Language models adapted on private text are often served through APIs, so privacy leakage occurs through generated outputs rather than exposed weights. Private prediction protects these releases. Methods such as PMixED incur privacy cost at each release and increasingly rely on the public model over long horizons. PAC privacy instead calibrates noise to output variability across possible secrets, adding less noise when predictions are stable. To our knowledge, PAC-private prediction has not previously been extended from classification to autoregressive generation. We construct $m=128$ overlapping worlds from the private corpus, with each record appearing in exactly $m/2$ worlds, and train one adapter per world over a frozen public model. The realized world is the secret. At each token, the public model defines a candidate set, the worlds vote, and their posterior-weighted disagreement determines the PAC noise; unanimity requires no calibration noise. We prove $I(S;Y_{1:T}) \leq I(S;H_T) \leq bT$. Our contributions are extending PAC privacy to autoregressive generation, handling adaptive self-generated contexts, and introducing coupled decoding that preserves privacy accounting while avoiding greedy degeneration. On WikiText-103 with GPT-2-small, we retain 74% of the fine-tuning gain at a per-token budget of $2^{-32}$, while membership-inference success is bounded by 51.08% after $10^6$ tokens; posterior-entropy estimates of leakage are roughly 17% of the charged budget. Inference privacy is not content protection: even when membership advantage on a memorized canary is indistinguishable from zero, the canary is emitted at the same rate. Against PMixED under matched membership-inference bounds on the same data universe and test set, we retain 98% of non-private headroom from $10^2$ to $10^6$ tokens, versus at most 56%, with no crossover.