生成式数据增强的可靠性:基于Wasserstein的理论与实证研究
On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study
浏览论文内容
中文总结 AI 辅助
本研究提出基于Wasserstein的条件生成式数据增强统计框架,推导泛化界,经实验发现增强可靠性由分布近似误差决定,为合成数据质量评估提供了新基础。
中文摘要 AI 辅助
生成式数据增强被广泛用于缓解类别不平衡问题,但其对下游泛化能力的理论影响仍鲜为人知。本研究开发了条件生成式数据增强的统计框架,并分析其对分类风险的影响。我们将增强形式化为分布混合过程,表明产生的风险失真由增强强度以及真实分布与生成分布间的类别条件Wasserstein距离共同控制。进一步基于Rademacher复杂度推导了依赖容量的泛化界,揭示了假设复杂度、增强强度与生成保真度间的明确权衡关系。实证层面,我们在二分类和多分类不平衡分类任务上,使用条件GAN(Conditional GAN)和条件WGAN-GP(Conditional WGAN-GP)增强对该框架进行评估。在所有数据集上,CWGAN-GP始终比CGAN实现更低的Wasserstein距离,表明其分布保真度更高。但更高的保真度并不一定转化为更优的分类性能,经典过采样方法往往仍具有竞争力。这些发现支持核心理论预测:增强的可靠性由分布近似误差而非仅预测性能决定。总体而言,本研究确立生成式数据增强为一种分布扰动过程,其可靠性可通过基于Wasserstein的度量量化,并由有限样本泛化保证支持。所提出的框架为仅基于分类准确率之外的合成数据质量评估提供了有原则的基础。
英文摘要
Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We formalize augmentation as a distribution-mixing process and show that the resulting risk distortion is controlled by both the augmentation strength and the class-conditional Wasserstein discrepancy between real and generated distributions. We further derive a capacity-dependent generalization bound based on Rademacher complexity, revealing an explicit trade-off between hypothesis complexity, augmentation intensity, and generative fidelity. Empirically, we evaluate the framework on binary and multiclass imbalanced classification tasks using Conditional GAN and Conditional WGAN-GP augmentation. Across datasets, CWGAN-GP consistently achieves lower Wasserstein discrepancies than CGAN, indicating improved distributional fidelity. However, improved fidelity does not necessarily translate into superior classification performance, with classical oversampling methods often remaining competitive. These findings support the central theoretical prediction that augmentation reliability is governed by distributional approximation error rather than predictive performance alone. Overall, this work establishes generative augmentation as a distributional perturbation process whose reliability can be quantified through Wasserstein-based measures and supported by finite-sample generalization guarantees. The proposed framework provides a principled foundation for evaluating synthetic data quality beyond classification accuracy alone.