StatD2GAN:当校准在合成天气序列的留出评估中掩盖生成器质量
StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences
- Hacettepe University(哈杰泰佩大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对天气序列生成模型评估中校准掩盖生成器质量的问题,本文提出留出协议重建StatD2GAN评估,证明校准使边际指标失效,并推荐下限参考评估与序列级方差分解。
中文摘要 AI 辅助
多变量天气序列的生成模型通常在边际校准后使用汇总分布指标进行常规评估。我们表明这种做法可能使架构结论失效,并围绕留出协议重建了StatD2GAN(一种具有进化权重调整的三判别器GAN)的评估:每个数据集的最后两个日历年作为留出集,并设置168小时禁令期,校准仅基于训练块进行拟合,所有指标均在留出块上计算。证据来自跨越五个柯本-盖革气候区的25个匹配(位置,种子)对,采用Holm校正下的Wilcoxon符号秩检验。得到四个结果。首先,等渗校准将Kolmogorov-Smirnov距离驱动到每个测试架构(包括一个故意弱化的RCGAN基线)的每位置噪声加偏移下限的2%以内,因此校准后的边际指标无法区分架构。其次,排序表示判别器是唯一移除后会显著降低跨变量依赖性(Kendall tau MAE +0.080,Holm p = 0.009)的组件,且具有依赖机制的效果:在安卡拉接近零,在迪拜和雅库茨克超过115%。秩变换变体将该机制隔离为对边际的分位数监督而非copula匹配。第三,物理约束违反是由校准而非生成器注入的;投影以可忽略的成本消除它们(deltaKS <= 0.003)。第四,汇总指标掩盖了TimeGAN中序列间周平均变异性(季节和机制多样性的代理)的崩溃,只有序列级统计才能揭示这一点。我们建议将下限参考边际评估、匹配对检验和序列级方差分解作为校准生成流程的最低要求。
英文摘要
Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS <= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.