arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03180cs.LG

跨合成数据生成器族的便携式因果公平性

Portable Causal Fairness Across Synthetic Data Generator Families

  • University of Washington Tacoma(华盛顿大学塔科马分校)

机构由 AI 辅助整理,请以论文原文为准。

Steven Golob, Sikha Pentyala, Martine De Cock

AI总结:

本研究将三种公平性定义迁移至多类合成数据生成器,验证了其通用性,提出的因果扩散主干生成公平性最优的合成数据,且对保真度影响极小,添加隐私保证不损害公平性。

AI中文摘要:

当统计机构或监管机构发布合成数据以替代敏感记录时,会选择生成表格的生成器,并可调整该生成器以消除不公平路径。DECAF在一个非私有生成对抗网络(GAN)上实现了这一思路:三种公平性定义对应生成器因果图上的三组边切割。该机制是否属于DECAF或因果分解本身尚未得到验证。我们将这三种定义迁移至三个不相关族的九种生成器(基于边际分布的、GAN、扩散模型,每种都包含差分隐私变体),在三个形式化隐私保证级别下,对Adult和COMPAS数据集完成了超过2520次配对运行。该机制可迁移至所有生成器,我们提出的新型因果扩散主干在所有测试族中生成了最公平的发布数据,且保真度接近边际分布层级。应用切割几乎不会改变保真度,仅使下游分类器的AUC平均下降0.07至0.15,且添加隐私保证不会降低数据的公平性。

英文摘要:

When a statistical agency or regulator releases synthetic data in place of sensitive records, it chooses the generator that produces the table, and can shape that generator so unfair pathways are absent. DECAF made this concrete on one non-private GAN: three fairness definitions become three sets of edge cuts on the generator's causal graph. Whether the mechanism belongs to DECAF, or to causal factorisation itself, was untested. We port all three definitions to nine generators from three unrelated families (marginals-based, GAN, and diffusion, each with differentially private variants), across three levels of formal privacy guarantee, over 2,520 matched-pair runs on Adult and COMPAS datasets. The mechanism transfers everywhere, and our new causal diffusion backbone yields the fairest release of any family we tested, at fidelity close to the marginals tier. Applying the cut barely moves fidelity, only costs a downstream classifier about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees don't make the data less fair.

↑