arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构化预测不确定性的最优传输Dropout

Optimal Transport Dropout for Structured Predictive Uncertainty

Giacomo Lorenzon, Francesco Regazzoni

arXiv 2609.33377首次发表:更新:

AI 中文总结

提出最优传输Dropout(OTD),通过可学习流传输潜在扰动以学习结构化预测不确定性,无需后验推断或集成,在合成和真实回归基准上优于蒙特卡洛Dropout。

AI 中文摘要

确定性神经网络和神经算子提供点预测,但缺乏内在的可靠性度量。然而,预测不确定性可能源于不可约的结果变异性、有限数据或所选模型类别的局限性。蒙特卡洛Dropout通过随机特征掩蔽提供了一种计算便捷的方式来构建预测分布,无需训练多个独立网络或显式推断模型参数的后验。然而,其扰动规律在很大程度上是预先设定的,且通常在各潜在坐标间因子化。我们引入了最优传输Dropout(OTD),该方法联合学习预测映射及其潜在扰动的规律。从一个简单的独立参考分布出发,OTD通过可学习的流传输潜在扰动,并将其传播通过预测神经网络,从而诱导出结构化的预测规律。训练使用严格适当的能量分数,而动力学作用项则对传输进行几何正则化。合成基准测试表明,OTD能够捕获多模态预测分布,在模型被错误指定时生成有意义的离散度,并在提供更多训练数据或更大模型容量时表现出收缩的离散度。对于场值偏微分方程代理模型,预测离散度与预测误差的空间模式强烈对齐。在此任务上,与蒙特卡洛Dropout相比,OTD产生更准确的预测和更好校准的、显著更窄的区间。在真实世界回归基准上,与已建立的基线相比,它进一步展示了有竞争力的准确性和更好的概率预测。因此,OTD提供了一种无需显式后验推断或独立训练预测器集成即可学习结构化预测不确定性的方法。

英文摘要

Deterministic neural networks and neural operators provide point predictions with no intrinsic measure of reliability. Yet, predictive uncertainty may stem from irreducible outcome variability, finite data, or limitations of the chosen model class. Monte Carlo dropout offers a computationally convenient way to construct a predictive distribution through stochastic feature masking, without training multiple independent networks or explicitly inferring a posterior over model parameters. However, its perturbation law is largely prescribed a priori and typically factorised across latent coordinates. We introduce Optimal Transport Dropout (OTD), which instead learns the predictive mapping and the law of its latent perturbations jointly. Starting from a simple independent reference distribution, OTD transports latent perturbations through a learnable flow and propagates them through the predictive neural network, thereby inducing a structured predictive law. Training uses the strictly proper Energy Score, while a kinetic-action term geometrically regularises the transport. Synthetic benchmarks show that OTD captures multimodal predictive distributions, generates meaningful dispersion when the model is misspecified, and exhibits contracting dispersion as more training data or greater model capacity are provided. For a field-valued partial differential equation surrogate, predictive dispersion strongly aligns with the spatial pattern of prediction errors. On this task, compared with Monte Carlo dropout, OTD yields more accurate predictions and better-calibrated, substantially narrower intervals. On real-world regression benchmarks, it further shows competitive accuracy and better probabilistic predictions compared to established baselines. OTD therefore offers a way to learn structured predictive uncertainty without explicit posterior inference or ensembles of independently trained predictors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑