arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DP-IPI:面向临床文本中间接个人标识符的混合差分隐私文本重写机制

DP-IPI: A Hybrid Differential Privacy Text Rewriting Mechanism for Indirect Personal Identifiers in Clinical Texts

Ibrahim Baroud, Stephen Meisenbacher, Sebastian Möller, Florian Matthes, Roland Roller

arXiv 2609.29684首次发表:更新:

发表机构

Technical University of Berlin; German Research Center for Artificial Intelligence (DFKI); Technical University of Munich; Munich Center for Machine Learning(柏林工业大学; 德国人工智能研究中心(DFKI); 慕尼黑工业大学; 慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DP-IPI,一种仅对临床文本中间接个人标识符片段进行噪声扰动的混合差分隐私重写方法,在降低重识别风险的同时保持文本质量与效用,实现更优的隐私-效用权衡。

AI 中文摘要

尽管现代匿名化和去标识化技术具有优势,但由于文本中残留的间接标识符,重识别的风险仍然显著。为解决这一问题,近期研究在差分隐私(DP)下应用文本重写,通过添加噪声扰动文本以防止数据链接。此类方法不加区分地私有化文本中的所有标记,降低了文本质量及其在临床等关键领域的可用性。聚焦于间接个人标识符(IPIs),我们引入了一种保持效用的DP文本重写方法,仅私有化包含IPIs的片段。我们证明,该方法能有效降低临床文本中的重识别风险,同时生成更连贯、更可用的输出文本,从而实现更高的隐私-效用权衡。在此,我们展示了混合文本私有化的有效性,它以高效、可用的方式发挥了DP的潜力。

英文摘要

Despite the strengths of modern anonymization and de-identification techniques, the risk of re-identification remains significant due to the indirect identifiers remaining in texts. To address this problem, recent works have applied text rewriting under Differential Privacy (DP) to prevent data linkage by perturbing texts via noise addition. Such methods privatize all tokens in a text indiscriminately, diminishing text quality and usability in critical domains such as in clinical settings. Focusing on indirect personal identifiers (IPIs), we introduce a utility-preserving DP text rewriting method that only privatizes spans containing IPIs. We show that our method effectively reduces re-identification risks in clinical texts while being producing more coherent and usable output texts, leading to higher privacy-utility trade-offs. In this, we demonstrate the effectiveness of hybrid text privatization, which leverages the promise of DP in an efficient, usable manner.

Comments15 pages, 5 figures, 5 tables, accepted to EMNLP 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑