AI 中文总结
针对噪声评估下的数据重塑问题,提出平坦共识扩散框架FCDiff,通过微观平滑与宏观共识聚合,在8数据集上取得最优稳健性。
AI 中文摘要
数据形状决定了特征如何结构化、模式如何分离以及分布如何覆盖底层领域。不良的数据形状可能使模型学习噪声而非可泛化的结构。本文研究以特征为中心的稳健数据重塑:在噪声评估和不完美数据条件下,生成仍然有用、稳定且可复现的特征变换。我们将重塑操作序列搜索视为奖励引导的扩散生成,并将稳健重塑视为在潜在奖励景观中搜索区域,而非孤立的奖励变换点。关键挑战是双重不稳定性:噪声评估器扭曲局部奖励引导,而随机生成轨迹可能收敛到不一致的解。我们提出FCDiff,一个平坦共识扩散框架,通过微观-宏观分解解决这两种失败。微观层用高斯平滑、蒙特卡洛平均梯度替代点估计奖励引导,将生成引导至局部平坦的奖励区域。宏观层用加权弗雷歇平均重心聚合独立引导的轨迹,选择共识支持的盆地并过滤随机异常值。在8数据集主队列中,在重尾评估器噪声下,FCDiff在下尾可靠性和稳健性方面取得了最佳总体排名,优于基于搜索的AutoFE和面向稳健性的生成基线,并在每个生成基线上具有统计显著的准确率提升。我们的结果表明,稳健的数据重塑需要搜索平坦、共识支持的区域,而非尖锐的单轨迹最优解。
英文摘要
Data shape determines how features are structured, how patterns are separated, and how distributions cover the underlying domain. Poor data shape can make models learn noise rather than generalizable structure. This paper studies robust feature-centric data reshaping: generating feature transformations that remain useful, stable, and reproducible under noisy evaluation and imperfect data conditions. We view reshaping operation sequence search as reward-guided diffusion generation, and robust reshaping as searching for regions in the latent reward landscape rather than isolated high-reward transformations. The key challenge is dual instability: noisy evaluators distort local reward guidance, while stochastic generative trajectories can converge to inconsistent solutions. We propose FCDiff, a flat-consensus diffusion framework that addresses both failures through a micro-macro decomposition. The micro layer replaces point-estimate reward guidance with Gaussian-smoothed, Monte Carlo averaged gradients, steering generation toward locally flat reward regions. The macro layer aggregates independently guided trajectories with a weighted Frechet-mean barycenter, selecting consensus-supported basins and filtering stochastic outliers. Across an 8-dataset headline cohort under heavy-tailed evaluator noise, FCDiff attains the best aggregate rank on lower-tail reliability and robustness against both search-based AutoFE and robustness-oriented generative baselines, with statistically significant accuracy gains over every generative baseline. Our results show that robust data reshaping requires searching for flat, consensus-supported regions rather than sharp single-trajectory optima.