发表机构
Arizona State University; Clemson University; NEC Laboratories America(亚利桑那州立大学; 克莱姆森大学; NEC美国实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对协变量偏移下数据增强的误导生成与结构不稳定问题,提出IGDPR框架,结合不变性引导扩散与原型重加权,提升数据质量与学习稳健性。
AI 中文摘要
在许多工业应用中,1)表格数据稀缺且不平衡,因此需要合成扩展;2)训练与部署之间的输入分布发生漂移(协变量偏移);3)验证集往往与未见过的测试环境存在差异;或4)标准生成模型仅仅模仿过时的源分布。这种学习设置限制了标准增强和自适应流程的稳定性。我们将此类设置下的任务泛化为“协变量偏移下的增强与加权学习”问题(AWL-CS)。AWL-CS对现有方法提出了两个关键挑战:1)误导性的生成引导,即模型优化的是源相似性而非下游任务相关性;2)分布密度的结构不稳定性,即重加权机制对噪声验证信号过拟合。为应对这些挑战,我们提出了IGDPR(不变性引导扩散与原型重加权),这是一个统一框架,协同了稳定合成与结构自适应:i)为实现任务相关的生成,我们利用不变性势引导扩散采样过程,确保合成样本与稳定的决策边界对齐,而非过时的相关性。ii)为实现稳定自适应,我们开发了一种基于原型的重加权策略,通过结构簇而非孤立点评估样本可靠性,有效过滤验证噪声。在真实数据上的大量实验表明,我们的方法通过增强对稳健学习最有益的数据来提高数据质量。
英文摘要
In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between training and deployment (covariate shift); 3) validation sets often diverge from unseen test environments; or 4) standard generative models simply mimic outdated source distributions. This learning setting limits the stability of standard augmentation and adaptation pipelines. We generalize the task under such setting as the Augmented and Weighted Learning under Covariate Shift problem (AWL-CS). AWL-CS imposes two critical challenges on existing methods: 1) misleading generative guidance where models optimize for source similarity rather than downstream task relevance, and 2) structural instability of distributional density where reweighting mechanisms overfit to noisy validation signals. To tackle these challenges, we propose IGDPR (Invariant-Guided Diffusion with Prototype Reweighting), a unified framework that synergizes stable synthesis and structural adaptation: i) To achieve task-relevant generation, we steer the diffusion sampling process using invariant potentials to ensure synthetic samples align with stable decision boundaries rather than outdated correlations. ii) To ensure stable adaptation, we develop a prototype-based reweighting strategy that assesses sample reliability through structural clusters instead of isolated points, effectively filtering validation noise. Extensive experiments on real data demonstrate our method improves data quality by augmenting the most beneficial data for robust learning.