发表机构
Johns Hopkins University; Indian Institute of Technology, Madras; Indian Institute of Science(约翰斯·霍普金斯大学; 马德拉斯印度理工学院; 印度科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对差分隐私语言模型存在的隐私与事实准确性的权衡问题,通过实验揭示差分隐私会增加模型幻觉风险,提出需开发兼顾隐私保障与事实准确性的干预措施。
AI 中文摘要
在医疗等关键领域,隐私和事实准确性都至关重要。令人担忧的是,我们发现并研究了差分隐私(DP)语言模型中的隐私-幻觉权衡问题。首先,我们通过实验表明,采用差分隐私预训练或微调的模型,相比非差分隐私模型会产生更多幻觉,且随着隐私预算变得更严格,幻觉的严重程度会增加。其次,我们研究了导致这种权衡的模型特性,证明差分隐私机制会使输出分布扁平化,可能将概率质量重新分配到事实错误的替代选项上。第三,通过控制训练数据中事实频率的实验,我们明确了信息频率可降低差分隐私模型的幻觉风险。总体而言,我们的研究结果强调需要更细致的隐私保护干预措施,既能提供严格的隐私保障,又不损害事实准确性。
英文摘要
Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre-trained or fine-tuned with DP tend to produce more hallucinations than non-DP counterparts, with increased severity as the privacy budget grows stricter. Second, we investigate model properties driving this tradeoff, demonstrating that DP mechanisms flatten output distributions, potentially redistributing probability mass toward factually incorrect alternatives. Third, through experiments where we control fact frequency in training data, we characterize how information frequency can reduce hallucination risks in DP models. Overall, our findings underscore the need for more nuanced privacy-preserving interventions that offer rigorous privacy guarantees without compromising factual accuracy.
CommentsAccepted to EMNLP 2026 (Findings)