在何处进行干预?对差分隐私合成表格数据上的公平感知学习进行基准测试
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
浏览论文内容
中文总结 AI 辅助
研究在差分隐私合成表格数据上公平干预的效果,以自适应迭代机制为基准,在多数据集、指标及策略下评估,比较四种管道配置,发现仅DP会降效,公平干预可部分恢复公平,后处理方法权衡更优,还开源相关内容以支持研究。
中文摘要 AI 辅助
机器学习模型在高风险领域的应用引发了对隐私和公平性的关注。差分隐私(DP)是隐私保护数据分析的黄金标准,而公平感知机制旨在减轻对代表性不足群体的歧视。但这两个目标可能冲突,DP常放大人口群体间的差异,且在DP约束下既定公平干预是否有效尚不清楚。本文首次对差分隐私合成表格数据上的公平干预进行系统评估。以自适应迭代机制(AIM)为基准,在四个数据集、多个群体公平性指标以及三类缓解策略(预处理、处理中、后处理)下,在广泛的隐私预算范围内评估公平干预。比较了四种管道配置:在原始数据上训练(基线);在DP合成数据上训练(仅DP);在原始数据上应用公平机制(仅公平);将公平机制与DP合成数据结合(DP+公平)。结果表明,仅DP会降低效用和公平性,应用公平干预可部分恢复公平结果。后处理方法在隐私预算和合成器之间往往能提供更稳定的公平-效用权衡,在保持竞争力效用的同时实现了显著的公平性提升。我们在开源存储库中发布了所有代码、数据和实验工件,以确保完全可重复性并支持未来关于隐私-公平性-效用权衡的研究。
英文摘要
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups. However, these objectives can conflict: DP often amplifies disparities across demographic groups, and little is known about whether established fairness interventions remain effective under DP constraints. In this work, we present, to our knowledge, the first systematic evaluation of fairness interventions on differentially private synthetic tabular data. Our benchmark centers on the Adaptive Iterative Mechanism (AIM), identified as the state-of-the-art marginal-based DP synthesizer (Cormode et al. 2025). We thus evaluate fairness interventions across four datasets, multiple group fairness metrics, and three categories of mitigation strategies (pre-processing, in-processing, and post-processing) under a wide range of privacy budgets. We compare four pipeline configurations: (Baseline) training on original data; (DP-only) training on DP synthetic data; (Fair-only) applying fairness mechanisms on original data; and (DP+Fair) combining fairness mechanisms with DP synthetic data. Our results demonstrate that while DP alone can degrade both utility and fairness, applying fairness interventions can partially restore equitable outcomes. Among them, post-processing methods tend to provide more stable fairness-utility trade-offs across privacy budgets and synthesizers, achieving strong fairness improvements while preserving competitive utility relative to other intervention stages. We release all code, data, and experimental artifacts in an open-source repository to ensure full reproducibility and to support future research on the privacy-fairness-utility trade-off.
发表机构
- ÉTS Montréal(蒙特利尔高等商学院)
- Inria Grenoble(格勒诺布尔计算机科学及自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。