发表机构
University of Calgary(卡尔加里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统评估15种机器学习模型,发现TabPFN在野火后泥石流预测中表现最优,合成数据增强可提升多数模型性能,结合可解释分析为该预测提供了全面框架。
AI 中文摘要
野火后泥石流的预测对于减轻近期 burned 区域强降雨期间对社区、基础设施和资源的危害至关重要。然而,由于特征空间中泥石流与非泥石流事件重叠、模型可解释性需求以及训练数据有限,识别可靠的机器学习模型变得复杂。本文通过从预测性能、特征重要性和合成数据增强三个方面对机器学习模型进行系统评估来应对这些挑战。利用美国西部野火后泥石流事件的流域尺度观测数据,我们比较了15种模型,包括表格先验数据拟合网络(TabPFN)。重复分层交叉验证显示,TabPFN在未增强数据下的性能最高,威胁评分为0.637,紧随其后的是最佳的基于树的模型。使用SHapley加性解释(SHAP)识别驱动预测的特征,发现短历时降雨强度和风暴累积量始终排名最高,而 burn severity 和地形特征贡献较小。我们进一步评估了使用TabPFN生成的样本进行合成数据增强,以解决泥石流观测数据稀缺的问题。合成数据增强提高了除卷积神经网络(CNN)外所有模型的性能,在深度学习模型中平均威胁评分增幅最大,为+0.041。通过结合严格的模型基准测试、可解释的特征分析和合成数据增强,本研究为改进野火后泥石流预测提供了全面框架。
英文摘要
Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas. However, identifying reliable machine learning models is complicated by overlapping debris-flow and non-debris-flow events in feature space, the need for model interpretability, and limited training data. This paper addresses these challenges through a systematic evaluation of machine learning models in terms of predictive performance, feature importance, and synthetic data augmentation. Using basin-scale observations of post-wildfire debris-flow events across the western United States, we compare 15 models, including the Tabular Prior-Data Fitted Network (TabPFN). Repeated stratified cross-validation shows that TabPFN achieves the highest unaugmented performance with a threat score of 0.637, closely followed by the best tree-based models. SHapley Additive exPlanations (SHAP) are used to identify the features driving predictions, revealing that short-duration rainfall intensity and storm accumulation consistently rank highest, while burn severity and terrain features contribute less. We further evaluate synthetic data augmentation using TabPFN-generated samples to address the scarcity of debris-flow observations. Synthetic augmentation improves the performance of all models except CNN, with the largest mean threat score increase of +0.041 among the deep learning models. By combining rigorous model benchmarking, interpretable feature analysis, and synthetic data augmentation, this work provides a comprehensive framework for improving post-wildfire debris-flow prediction.