arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WildfireSpreadBench:在野火蔓延预测中,指标决定模型

WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction

Arin Gopakumar, Marco Pannozzo

arXiv 2609.22191首次发表:更新:

发表机构

UC Berkeley; Purdue University(加州大学伯克利分校; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在野火蔓延预测中对比多种模型,发现仅用平均精度(AP)评估会偏向预测面积过大或过小的模型,而阈值相关指标更能反映实际可用性,并提出三种预测特征以指导模型选择。

AI 中文摘要

机器学习正越来越多地被用于预测活跃野火次日将燃烧的区域,从而为疏散边界和火线控制提供信息。大多数模型使用平均精度(AP)进行评估,该指标汇总了所有决策阈值下的性能,尽管基于预测采取行动需要选择一个阈值。我们使用共享的评估流程和两种输入配置,在WildfireSpreadTS上对五种判别式架构和一种生成模型进行了基准测试。我们发现,模型排名会因性能是通过AP衡量还是通过F1和IoU等阈值相关指标衡量而有所不同。AP最高的模型标记的面积是实际燃烧面积的4至5倍,在F1和IoU上排名六分之五,而召回率最高的模型标记的面积是实际燃烧面积的16至23倍。具有更可用预测的模型的AP分数低24至37。在不同架构中,我们识别出三种不同的预测特征:过度预测、平衡和欠预测,仅凭AP无法区分这些特征。将输入从7个通道扩展到23个通道,AP平均变化0.03,而架构间的差异为0.21至0.24。这些结果表明,仅使用AP可能会偏向那些预测不适合业务化野火预报的模型。

英文摘要

Machine learning is being increasingly used to predict where active wildfires will burn the following day, helping inform evacuation boundaries and containment lines. Most models are evaluated using Average Precision (AP), which summarizes performance across all decision thresholds, although acting on a forecast requires choosing one. We benchmarked five discriminative architectures and one generative model on WildfireSpreadTS using a shared evaluation pipeline and two input configurations. We found that model rankings varied depending on whether performance was measured by AP or by threshold-dependent metrics like F1 and IoU. The highest-AP model flagged 4 to 5 times the area that burned and ranked fifth of six on F1 and IoU, and the most recall-heavy model flagged 16 to 23 times. Models with more usable predictions had AP scores 24 to 37 lower. Across architectures, we identified three distinct prediction profiles: over-predicting, balanced, and under-predicting, which AP alone could not distinguish. Expanding the input from 7 to 23 channels changed AP by 0.03 on average, against a 0.21 to 0.24 spread across architectures. These results show AP alone can favor models whose predictions are poorly suited for operational wildfire forecasting.

Comments11 pages, 2 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑