发表机构
National Chung Hsing University; Texas Instruments; Academia Sinica; National Tsing Hua University(国立中兴大学; 德州仪器; 中央研究院; 国立清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于样本的生成模型预测分布存在偏差和校准不良的问题,提出偏差校正的保形PIT校准框架,输出校准的预测分布,实现有限样本校准和嵌套区间,并在模拟和降水预报中显著提升性能。
AI 中文摘要
条件生成模型,包括扩散模型和集成预报器,常常生成预测样本,但缺乏可处理的似然表示。这种基于样本的预测分布可能系统性地存在偏差且校准不良。我们提出偏差校正的保形概率积分变换(PIT)校准,这是一种分割样本的后处理框架,输出校准的预测分布而非单一固定水平的预测区间。该方法首先在保留的偏差分割上估计仿射位置-尺度校正,然后使用保形校准器校准随机化PIT值。所得的预测律表示为生成器顺序统计量上的加权经验分布,从而能够直接计算阈值一致的超越概率、任意分位数、最高密度区间、预期尾部损失和校准的重采样。相比之下,标准保形预测主要提供固定水平的预测集或阈值决策,并不直接估计预测概率或高密度区域。我们建立了在可交换性下有限样本概率校准,并展示了基于PIT中心性得分的可选分割保形包装器,可在用户指定水平下提供具有有限样本边际覆盖的嵌套预测区间。带有受控误设的模拟研究和WeatherBench-2降水预报应用表明,相对于未校准的基于样本的预测和仅区间的保形基线,概率校准和下游分布摘要均有显著改进。
英文摘要
Conditional generative models, including diffusion models and ensemble forecasters, often produce predictive samples without a tractable likelihood representation. Such sample-based predictive distributions can be systematically biased and poorly calibrated. We propose bias-corrected conformal probability integral transform (PIT) calibration, a split-sample post-processing framework that outputs a calibrated predictive distribution rather than a single fixed-level prediction interval. The method first estimates an affine location-scale correction on a held-out bias split, then calibrates randomized PIT values using a conformal calibrator. The resulting predictive law is represented as a weighted empirical distribution on the generator order statistics, enabling the direct computation of threshold-coherent exceedance probabilities, arbitrary quantiles, highest-density intervals, expected tail losses, and calibrated resamples. In contrast, standard conformal prediction primarily provides fixed-level prediction sets or threshold decisions and does not directly estimate predictive probabilities or high-density regions. We establish finite-sample calibration in probability under exchangeability and show how an optional split-conformal wrapper based on a PIT-centrality score gives nested prediction intervals with finite-sample marginal coverage at user-specified levels. Simulation studies with controlled misspecification and a WeatherBench-2 precipitation-forecasting application demonstrate substantial improvements in probabilistic calibration and downstream distributional summaries relative to uncalibrated sample-based forecasts and interval-only conformal baselines.