arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11335cs.CLcs.AI

匿名化对大型语言模型性能的影响

On the Impact of Anonymization on the Performance of Large Language Models

Tobias Deußer, Max Hahnbück, Lorenz Sparrenberg, Tobias Uelwer, Christian Bauckhage, Rafet Sifa

首次发表
浏览论文内容

中文总结 AI 辅助

本研究系统评估了匿名化对五个大型语言模型在十一个基准上的性能影响,发现性能下降因模型能力和任务类型而异,并建议匿名化需与模型和任务协同设计以平衡隐私与效用。

中文摘要 AI 辅助

随着大型语言模型越来越多地部署在敏感领域,对输入数据进行匿名化以保护个人身份信息已成为一种关键实践。然而,这种匿名化对模型效用的影响尚未得到充分理解。本文对隐私与性能之间的权衡进行了系统的实证研究。我们评估了五个主流语言模型在十一个多样化基准上的表现,比较了它们在原始输入与假名化输入上的性能。我们的结果显示,虽然匿名化通常会降低性能,但影响高度细微。我们发现,能力更强的模型,如Qwen2.5-72B和GPT-4o mini,遭受的性能下降最大,这表明它们对特定实体信息的依赖更强。影响也依赖于任务:TruthfulQA上的性能在匿名化后有所提升,而像RGB这样的检索聚焦任务则经历了灾难性的下降。进一步的实验表明,保留实体唯一性的可逆匿名化技术显著优于不可逆技术(如删除),而明确提示模型关于匿名化并未带来明显益处。我们得出结论,匿名化并非一刀切的解决方案,必须与模型和任务共同设计,以有效平衡隐私与效用。我们的发现为开发更稳健、隐私感知的AI系统提供了关键基线。

英文摘要

As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper presents a systematic empirical study of the trade-off between privacy and performance. We evaluate five prominent language models across eleven diverse benchmarks, comparing their performance on original versus pseudonymized inputs. Our results reveal that while anonymization generally degrades performance, the effect is highly nuanced. We find that more capable models, such as Qwen2.5-72B and GPT-4o mini, suffer the largest performance drops, suggesting a stronger reliance on specific entity information. The impact is also task-dependent: performance on TruthfulQA improves with anonymization, while retrieval-focused tasks like RGB experience a catastrophic decline. Further experiments show that reversible anonymization techniques that preserve entity uniqueness significantly outperform irreversible ones like redaction, and that explicitly prompting models about anonymization offers no discernible benefit. We conclude that anonymization is not a one-size-fits-all solution and must be co-designed with the model and task in mind to balance privacy and utility effectively. Our findings provide a crucial baseline for developing more robust, privacy-aware AI systems.

发表机构

  • University of Bonn(波恩大学)
  • Fraunhofer IAIS(弗劳恩霍夫智能分析与信息系统研究所)
  • Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习和人工智能研究所)
  • Microsoft Germany GmbH(微软德国有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑