arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于审计生成式医学影像中隐私与公平性的数据干预框架

A Data-Interventional Framework for Auditing Privacy and Fairness in Generative Medical Imaging

Mischa Dombrowski, Bernhard Kainz

arXiv 2609.26623首次发表:更新:

发表机构

Friedrich-Alexander-Universität Erlangen-Nürnberg; Imperial College London(埃尔朗根-纽伦堡大学; 伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出数据干预框架,利用合成解剖指纹作为探针,系统分析扩散模型在医学影像生成中的隐私与公平性张力,发现模型对稀有特征要么遗忘要么记忆,并引入指标t'揭示条件稀有性与记忆化的关系,为安全公平的数据共享提供策略。

AI 中文摘要

基于扩散的合成数据生成为共享医学影像数据而不泄露敏感患者记录提供了一条有前景的途径。然而,生成模型在隐私与公平性之间面临根本性张力:它们可能记忆稀有训练样本,导致隐私风险,或无法再现代表性不足的特征,导致不公平的合成分布。虽然先前的工作大多孤立地关注记忆化或公平性,但它们的相互作用仍未被充分理解。在本工作中,我们引入了一个数据干预框架,以系统分析扩散模型中的隐私与公平性。我们讨论了合成解剖指纹(SAFs),即稀有且手动注入的图像特征,作为受控探针,研究模型是否跨身份泛化敏感属性、记忆训练样本,或完全抑制稀有信号。在多种条件模态中,我们观察到一致的行为:模型要么遗忘这些指纹,要么记忆包含它们的整个图像,但不会将它们泛化到新图像。为了支持显式样本提取不可行的大规模审计,我们进一步引入了指标t',该指标利用扩散过程的内部结构来估计模型对记忆化的敏感性。通过比较不同惊奇度的条件信号,我们揭示了条件稀有性与记忆化行为之间的明确关系。高惊奇度的条件信号作为检索键,放大记忆化,而低惊奇度的条件信号则系统性地抑制稀有特征,即使这些特征在训练数据中反复出现。我们的发现为安全且公平的合成医学数据共享提供了可操作的见解和具体的缓解策略。代码可在以下网址获取:此 https URL。

英文摘要

Diffusion-based synthetic data generation offers a promising route for sharing medical imaging data without releasing sensitive patient records. However, generative models face a fundamental tension between privacy and fairness: they may memorize rare training samples, leading to privacy risks, or fail to reproduce underrepresented features, resulting in unfair synthetic distributions. While prior work has largely focused on either memorization or fairness in isolation, their interaction remains insufficiently understood. In this work, we introduce a data-interventional framework to systematically analyze privacy and fairness in diffusion models. We discuss synthetic anatomical fingerprints (SAFs), rare and manually injected image features, as controlled probes to study whether models generalize sensitive attributes across identities, memorize training samples, or suppress rare signals entirely. Across multiple conditioning modalities, we observe a consistent behavior: models either forget these fingerprints or memorize the entire image in which they appear, but do not generalize them to novel images. To support large-scale auditing where explicit sample extraction is infeasible, we further introduce the indicator metric t', which estimates a model's susceptibility to memorization by exploiting the internal structure of the diffusion process. By comparing conditioning signals of varying surprisal, we reveal a clear relationship between conditioning rarity and memorization behavior. Highly surprising conditioning signals act as retrieval keys that amplify memorization, whereas low-surprisal conditioning signals systematically suppress rare features, even when these appear repeatedly in the training data. Our findings provide actionable insights and concrete mitigation strategies for safe and fair synthetic medical data sharing. Code is available at https://github.com/MischaD/Privacy.

CommentsAccepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://melba-journal.org/2026:026

Journal refMachine.Learning.for.Biomedical.Imaging. 2026 (2026)

DOI:10.59275/j.melba.2026-92a6

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑