无遗憾的隐私:差分隐私推理时对齐
Privacy Without Regret: Differentially Private Inference-Time Alignment
浏览论文内容
中文总结 AI 辅助
该研究针对最佳N采样的奖励黑客攻击与偏好数据隐私问题,提出PrivBoN和PrivITP方法,实现差分隐私推理时对齐,经实验验证其性能优于原有策略。
中文摘要 AI 辅助
最佳N(BoN)采样是最简单且应用最广泛的推理时对齐策略,但它存在两个不同的问题:奖励黑客攻击,即所选响应利用代理奖励模型中的错误;以及用于训练该奖励模型的敏感人类偏好数据缺乏任何隐私保护。我们表明,单一干预措施——在选择前向奖励分数添加校准噪声——可同时解决这两个问题。我们的第一个结果,私有最佳N(PrivBoN),确立了适当尺度的Gumbel噪声可同时提供ε-差分隐私并实现KL正则化对齐。每当隐私预算超过临界阈值ε*时,隐私要求的噪声是遗憾最优正则化,且隐私施加的额外对齐成本为零,匹配Huang等人(2025)的信息论上限。由于ε*依赖于未知的覆盖系数,我们引入私有推理时悲观主义(PrivITP),它结合χ²正则化拒绝采样与两阶段高斯机制。PrivITP实现事后(ε,δ)-差分隐私,其隐私成本独立于响应数量n,将正则化参数与隐私参数明确解耦,且在噪声膨胀项内达到上限。在多个语言模型、数据集和奖励模型上的实验证实了我们的结果:PrivBoN和PrivITP是规模单调的(不同于BoN,其在超过临界n后性能下降),且在同等隐私水平下PrivITP与PrivBoN相当或优于PrivBoN,在强隐私机制中收益最大。
英文摘要
Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in which the selected response exploits errors in the proxy reward model, and the absence of any privacy protection for the sensitive human preference data used to train that reward model. We show that a single intervention-adding calibrated noise to reward scores before selection-resolves both. Our first result, Private Best-of-N (PrivBoN), establishes that Gumbel noise at an appropriate scale simultaneously provides $ε$-differential privacy and implements KL-regularized alignment. Whenever the privacy budget exceeds a critical threshold $ε^*$, the privacy-mandated noise is the regret-optimal regularization, and privacy imposes zero additional alignment cost-matching the information-theoretic skyline of Huang et al. (2025). Because $ε^*$ depends on an unknown coverage coefficient, we introduce Private Inference-Time Pessimism (PrivITP), which combines $χ^2$-regularized rejection sampling with a two-phase Gaussian mechanism. PrivITP achieves ex-post $(ε,δ)$-DP with a privacy cost independent of the number of responses $n$, cleanly decouples the regularization parameter from the privacy parameter, and attains the skyline up to a noise-inflation term. Experiments across several language models, datasets, and reward models confirm our results: PrivBoN and PrivITP are scaling-monotonic (unlike BoN, which degrades past a critical $n$), and PrivITP matches or outperforms PrivBoN at equivalent privacy levels, with the largest gains in the strong-privacy regime.
发表机构
- Indian Institute of Technology Kanpur(坎普尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。