AI 中文总结
该研究针对DPO的隐私缺陷,提出PrivDPO方法,通过沿偏好轴添加校准随机性实现偏好隐私,在对齐基准和LLM上取得了更好的隐私-效用权衡。
AI 中文摘要
直接偏好优化(DPO)现已成为使用人类偏好数据对齐大语言模型(LLM)的标准方法。每个DPO示例包含一个提示和一对候选模型响应。虽然提示和响应通常是公开的或模型生成的,但响应之间的相对偏好反映了主观判断,可能会暴露标注者或最终用户的敏感属性。现成的隐私保护方法与该结构匹配度不佳,会导致训练中注入不必要的噪声和有偏更新。本文将偏好隐私形式化为一种针对DPO的标签差分隐私(label-DP)式隐私概念,假设攻击者已知道提示和响应,仅保护候选响应之间的相对偏好。随后设计了PrivDPO,这是一种DPO变体,可在保持与大规模LLM训练兼容性的同时执行偏好隐私。核心发现是,对于仅在偏好信号上不同的相邻示例,梯度差异位于仅由文本确定的一维偏好轴上;所有偏好信息均通过该轴流动。PrivDPO通过对DPO目标进行无偏随机重缩放,仅沿该轴添加校准后的随机性,避免了逐示例梯度操作。在三个对齐基准和三个LLM家族上的实验表明,与隐私保护基线相比,PrivDPO始终实现了出色的隐私-效用权衡。
英文摘要
Direct preference optimization (DPO) is now a standard method for aligning large language models (LLMs) using human preference data. Each DPO example contains a prompt and a pair of candidate model responses. While prompts and responses are often public or model-generated, the relative preference between responses reflects subjective judgments and can reveal sensitive attributes of annotators or end users. Off-the-shelf privacy-preserving approaches are not well matched to this structure, leading to unnecessary noise injection and biased updates in training. In this paper, we formalize preference privacy, a label-DP-style privacy notion for DPO that protects only the relative preference between candidate responses, assuming an adversary who already knows the prompt and responses. We then design PrivDPO, a DPO variant that enforces preference privacy while remaining compatible with large-scale LLM training. Our main observation is that, for neighboring examples differing only in their preference signal, the gradient difference lies on a one-dimensional preference axis determined solely by the text; all preference information flows through this axis. PrivDPO adds calibrated randomness only along this axis via an unbiased randomized rescaling of the DPO objective, avoiding per-example gradient operations. Our experiments on three alignment benchmarks and three LLM families show that PrivDPO consistently achieves strong privacy-utility trade-offs compared with privacy-preserving baselines.
Commentsaccepted for publication at CCS 2026, extended version