arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过偏好优化实现大语言模型对齐中的隐私、效用与安全平衡

Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization

Dishu Yang, Jingjing Liu, Jize Li

arXiv 2608.30141首次发表:更新:

AI 中文总结

该研究提出P3M协议,在不修改目标函数或引入正式隐私机制的情况下,通过调整隐私-偏好数据比例,降低LLM的金丝雀记忆与成员推理攻击性能,同时平衡隐私、效用与安全。

AI 中文摘要

偏好优化被广泛用于使大语言模型与人类偏好对齐,但偏好数据的组成也可能影响与隐私相关的记忆。本文研究在不修改目标函数或引入正式隐私机制的情况下,向直接偏好优化(DPO)中添加合成隐私-偏好对是否与基于金丝雀的记忆信号降低相关。我们提出隐私压力偏好混合(P3M),这是一种数据组成协议,在保持有用性和无害性偏好数据固定的同时,改变隐私-偏好数据的数量。我们使用Gemma 3 270M-IT在5个随机种子下评估非隐私基线以及0.5、1.0和2.0的隐私混合比例,并使用4位量化的Gemma 2 2B-IT在3个种子下验证相同的4种条件。总体而言,在所测试的条件下,隐私-偏好混合与两种模型设置下的平均金丝雀后缀对数似然代理值降低相关,且在混合源2B评估中,与基线相比,聚合成员推理攻击性能降低。具体而言,在隐私感知2B配置中,平均受试者工作特征曲线下面积(AUROC)为0.596至0.629,平均精确率-召回率曲线下面积(AUPRC)为0.541至0.575,而基线分别为0.804和0.790。不过,成员区分度的降低并非在所有数据源中都一致成立。此外,隐私比例与无害性偏好准确率之间的关系因模型设置而异,而有用性偏好准确率总体保持稳定。这些发现表明,P3M应被视为一种用于研究隐私-效用-安全权衡的轻量级经验协议,而非正式隐私保证或针对提取攻击的防御手段。

英文摘要

Preference optimization is widely used to align large language models with human preferences, but preference-data composition may also influence privacy-relevant memorization. We examine whether adding synthetic privacy-preference pairs to Direct Preference Optimization (DPO) is associated with lower canary-based memorization signals without modifying the objective or introducing a formal privacy mechanism. We propose Privacy-Pressure Preference Mixing (P3M), a data-composition protocol that varies the amount of privacy-preference data while keeping helpfulness and harmlessness preference data fixed. We evaluate a non-privacy Baseline and privacy-mixing ratios of 0.5, 1.0, and 2.0 using Gemma 3 270M-IT across five random seeds and validate the same four conditions using 4-bit-quantized Gemma 2 2B-IT across three seeds. Overall, under the tested conditions, privacy-preference mixing is associated with lower mean canary suffix log-likelihood proxy values across both model settings and lower aggregate membership-inference attack performance relative to the Baseline in the mixed-source 2B evaluation. Specifically, across the privacy-aware 2B configurations, the mean area under the receiver operating characteristic curve (AUROC) ranges from 0.596 to 0.629, and the mean area under the precision-recall curve (AUPRC) ranges from 0.541 to 0.575, compared with 0.804 and 0.790, respectively, for the Baseline. However, the reduction in membership distinguishability does not hold uniformly across data sources. Moreover, the relationship between the privacy ratio and harmlessness preference accuracy varies by model setting, whereas helpfulness preference accuracy remains broadly stable. These findings suggest that P3M should be viewed as a lightweight empirical protocol for examining privacy-utility-safety trade-offs rather than as a formal privacy guarantee or a defense against extraction attacks.

Comments6 pages, accepted for presentation at PRAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑