arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多层次回归树及其在美国西部野火中的应用

Multilevel regression trees with application to wildfires in the American west

John Henry V. Gray, Tianjian Zhou, Benjamin A. Shaby

arXiv 2609.36328首次发表:更新:

发表机构

Colorado State University; Takeda Pharmaceuticals(科罗拉多州立大学; 武田制药)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出多层次贝叶斯回归树模型,通过生态区间信息共享和并行回火算法,提升美国西部野火预测性能并揭示关键驱动变量。

AI 中文摘要

我们提出了一种在多层次结构内拟合的贝叶斯回归树模型,并将其应用于美国西部的历史野火数据。相关群体(生态区)之间的信息共享与高度可解释的回归树相结合,使得我们能够更好地预测和理解与野火相关的气候和土地覆盖变量。通过使用一系列性能指标进行模拟研究,我们证明了我们的方法产生的树后验分布在结构上与假设的真实树最为相似,同时实现了良好的样本外预测性能。应用于大型野火数据集时,我们详细探讨了与每个生态区对应的回归树内的变量分裂,并考虑了每个地点的已知特征。树之间的共享超参数提供了对所有生态区预测野火中变量和分裂值重要性的高度有用理解,这在可比模型中无直接对应。具体而言,我们强调潜在蒸发量、温度和常绿森林土地覆盖是与历史野火最相关的变量,并且在各组中最常选择的分裂值中观察到一些可辨别的模式。我们提出了一种基于并行回火的新算法,在真实后验温度下以共享超参数为条件,改善了马尔可夫链混合,这是贝叶斯CART模型中的一个已知瓶颈。

英文摘要

We propose a Bayesian regression tree model fit within a multilevel structure and apply it to historic wildfire data in the western United States. Sharing of information between related groups (ecoregions) combined with highly interpretable regression trees allows for better predictions and understanding of climate and land cover variables predictive of wildfires. By doing a simulation study with a range of performance metrics, we demonstrate our method produces tree posteriors most structurally similar to assumed true trees, while simultaneously achieving good out-of-sample predictive performance. Applied to a large wildfire data set, we explore variable splits within regression trees corresponding to each ecoregion in detail, taking into account known features of each location. Shared hyperparameters between trees provide highly useful understanding of both variable and split value importance in predicting wildfires among all ecoregions, with no direct parallel in comparable models. Namely, we highlight potential evaporation, temperature, and evergreen forest land cover as variables most associated with historic wildfires, with some observable patterns in split values most commonly chosen across the groups. We propose a new algorithm based on parallel tempering, conditioning on shared hyperparameters at the true posterior temperature, improving Markov chain mixing, a known bottleneck in Bayesian CART models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑