arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28945cs.LG

FairDiffuseVQVAE:基于向量量化隐变量条件优化的表格扩散模型采样时刻公平性

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents

Nitish Nagesh, Mahdi Bagheri, Amir M. Rahmani

首次发表
浏览论文内容

中文总结 AI 辅助

FairDiffuseVQVAE是将保真度与公平性解耦的两阶段表格扩散模型,在多个数据集上实现了显著的公平性提升,同时保持了优异的样本质量。

中文摘要 AI 辅助

合成表格数据正日益用于隐私保护数据共享、数据增强以及缓解下游分类器偏差。当前最先进的表格扩散模型如TabDDPM和TabSyn具备出色的分布保真度,但不提供任何公平性机制;相反,感知公平性的表格生成器(DECAF、FairTGAN、FairTabDDPM)会在训练时施加显式公平性惩罚,虽能获得一定公平性提升,但会大幅牺牲样本质量或下游效用。本文提出FairDiffuseVQVAE,这是一种将保真度与公平性解耦的两阶段架构:第一阶段是带有行级判别器的向量量化自编码器(无公平性项),第二阶段是采用无分类器引导的DiffuseVAE式连续扩散优化器,其以第一阶段的重构结果和受保护属性为条件。公平性作为采样分布的属性自然产生——推理时对受保护属性的均匀采样可通过构造实现人口 parity,而非来自竞争损失项。在Adult、Bank和COMPAS数据集上,FairDiffuseVQVAE取得最高的平均人口 parity比率(0.702,较FairTabDDPM提升47%)和平等机会比率(0.686,提升100%);同时,它还达到了所有已发表方法中最低的平均成对相关误差(0.034),同时明确用约15个AUC点换取这些公平性提升。

英文摘要

Synthetic tabular data is increasingly used in privacy-preserving data sharing, data augmentation, and to mitigate downstream classifier bias. State-of-the-art tabular diffusion models such as TabDDPM and TabSyn achieve excellent distributional fidelity but offer no mechanism for fairness; conversely, fairness-aware tabular generators (DECAF, FairTGAN, FairTabDDPM) impose explicit fairness penalties at training time, yielding modest fairness gains at substantial cost to either sample quality or downstream utility. We introduce FairDiffuseVQVAE, a two-stage architecture that decouples fidelity from fairness: a vector-quantized autoencoder with a row-level discriminator (Stage~1, no fairness terms) is followed by a DiffuseVAE-style continuous diffusion refiner that conditions on both the Stage-1 reconstruction and the protected attribute via classifier-free guidance (Stage~2). Fairness emerges as a property of the sampling distribution -- uniform sampling of the protected attribute at inference time enforces demographic parity by construction, rather than from competing loss terms. On the Adult, Bank and COMPAS datasets, FairDiffuseVQVAE achieves the highest mean Demographic Parity Ratio ($0.702$, $+47\%$ over FairTabDDPM) and Equalized Odds Ratio ($0.686$, $+100\%$). It also attains the lowest mean pair-wise correlation error ($0.034$) of any published method, while explicitly trading $\sim$$15$ AUC points for these fairness gains.

发表机构

  • University of California, Irvine(加州大学欧文分校)

机构由 AI 辅助整理,请以论文原文为准。

↑