arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

比较统计学习模型在污水流行病学中的应用:以诺如病毒为例

Comparing statistical learning models in wastewater-based epidemiology: An application to norovirus

Caelan McNamara, Fion Tan, Ella White, Elizaveta Semenova, Marta Blangiardo

arXiv 2609.20038首次发表:更新:

发表机构

School of Public Health, Imperial College London; MRC Centre for Environment and Health, Imperial College London; MRC Centre for Global Infectious Disease Analysis, Imperial College London(帝国理工学院公共卫生学院; 帝国理工学院环境与健康医学研究委员会中心; 帝国理工学院全球传染病分析医学研究委员会中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究比较六种统计学习模型预测污水中诺如病毒浓度,发现随机森林综合性能最优,INLA-SPDE贝叶斯模型在不确定性量化上表现佳,并揭示预测精度与计算效率的权衡。

AI 中文摘要

污水流行病学(WBE)是传染病监测中日益重要的工具,但针对预测病原体浓度在空间和时间上的建模方法,目前缺乏直接的比较研究。我们以英格兰的诺如病毒为案例,比较了六种建模方法的预测性能,考虑了2021年5月至2022年3月期间作为英国健康保护环境监测计划一部分收集的来自152个污水处理厂的3,232份污水样本。基准模型是使用集成嵌套拉普拉斯近似(INLA)结合随机偏微分方程方法(SPDE)的贝叶斯时空模型,并与Lasso回归、广义加性模型(GAM)、贝叶斯GAM、极端梯度提升(XGBoost)和随机森林进行比较。模型通过10折空间块交叉验证进行评估,评估指标包括均方误差、偏差、相关性、95%预测区间的经验覆盖率、区间得分和计算成本。随机森林是整体表现最佳的模型,取得了最好的区间得分和最接近名义经验覆盖率(95.3%),同时保持了与XGBoost相当的点预测精度。INLA-SPDE模型也表现良好,具有接近名义的经验覆盖率、第三好的区间得分,并且在所有评估指标中持续保持低偏差。比较预测的时空趋势,两种模型产生了相似的空间模式,但在不确定性估计方面存在显著差异。我们的研究结果突显了预测精度、不确定性量化和计算效率之间的权衡。集成机器学习方法非常适合快速预测,而贝叶斯地质统计模型在公共卫生监测中优先需要概率决策支持时仍然具有价值。

英文摘要

Wastewater-based epidemiology (WBE) is an increasingly important tool for infectious disease surveillance, but there has been limited direct comparison of modelling approaches for predicting pathogen concentrations across space and time. We compare the predictive performance of six modelling approaches using norovirus in England as a case study, considering 3,232 wastewater samples from 152 sewage treatment works collected between May 2021 and March 2022 as part of the UK Environmental Monitoring for Health Protection programme. The benchmark was a Bayesian spatio-temporal model using Integrated Nested Laplace Approximation (INLA) with the Stochastic Partial Differential Equation approach (SPDE), compared with Lasso regression, Generalised Additive Models (GAM), Bayesian GAM, Extreme Gradient Boosting (XGBoost), and Random Forest. Models were evaluated using 10-fold spatial-block cross-validation with metrics including mean squared error, bias, correlation, empirical coverage of 95% prediction intervals, interval score, and computational cost. Random Forest was the best overall performing model, achieving the best interval score and nearest to nominal empirical coverage (95.3%), while maintaining point prediction accuracy comparable to XGBoost. The INLA-SPDE model also performed well, with near-nominal empirical coverage, the third best interval score, and consistently low bias across all evaluated metrics. Comparing predicted spatio-temporal trends, both models yielded similar spatial patterns but notable differences in uncertainty estimation. Our findings highlight a trade-off between predictive accuracy, uncertainty quantification, and computational efficiency. Ensemble machine learning methods are well suited to rapid prediction, whereas Bayesian geostatistical models remain valuable when probabilistic decision support is a priority for public health surveillance

Comments25 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑