arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Reservoir:用于训练和评估流行病学模型的大规模模拟数据集

Reservoir: A Large-Scale Simulated Dataset for Training and Evaluating Epidemiological Models

Carson Dudley, Reiden Magdaleno, Marisa Eisenberg

arXiv 2608.27408首次发表:更新:

AI 中文总结

本文介绍Reservoir,含50万条暴发轨迹的大规模流行病模拟数据集,用于解决现实流行病学数据规模不足的问题,支持反事实实验等,可训练大规模流行病模型。

AI 中文摘要

大规模标准化数据集推动了从蛋白质结构预测到自然语言处理等AI驱动的科学建模领域的诸多进展。传染病流行病学正越来越多地采用AI方法进行预测、监测和暴发分析,但用于训练这些方法的时间序列数据规模仍比其他领域取得进展所依赖的语料库小几个数量级。由于现实世界流行病学数据的范围实际上无法达到训练真正大规模AI方法所需的规模,模拟数据提供了一种可行的替代方案。本文介绍Reservoir,这是一个大型开源模拟器及真实流行病模拟数据集,其中每条轨迹都带有完整的真实标签,包括现实暴发中无法直接测量的量,如真实感染数、随时间变化的再生数和反事实干预效果。Reservoir由具有真实噪声和报告伪影的随机模拟器生成,同时包含可配置时间、依从性和年龄依赖效力的干预措施。当前版本包含50万条暴发轨迹,覆盖10亿个模拟日,涉及多种病原体特征、人口结构和干预方案。Reservoir支持反事实实验、监测设计研究以及训练现实世界数据集无法提供的规模的流行病模型。

英文摘要

Large-scale, standardized datasets have driven many advances in AI-based scientific modeling, from protein structure prediction to natural language processing. Infectious disease epidemiology is increasingly adopting AI methods for forecasting, surveillance, and outbreak analytics, but the time-series data available to train them remains orders of magnitude smaller than the corpora behind the advances seen in other fields. Because the scope of real-world epidemiological data cannot practically reach the scale needed to train truly large-scale AI methods, simulated data provides a possible alternative. Here we introduce Reservoir, a large open simulator and dataset of realistic epidemic simulations in which every trajectory carries complete ground-truth labels, including quantities that cannot be measured directly in a real outbreak, such as true infection counts, time-varying reproduction numbers, and counterfactual intervention effects. Reservoir is generated by a stochastic simulator with realistic noise and reporting artifacts, together with interventions with configurable timing, compliance, and age-dependent efficacy. The current release contains 500,000 outbreak trajectories spanning one billion simulated days across diverse pathogen characteristics, population structures, and intervention regimes. Reservoir enables counterfactual experiments, surveillance-design studies, and training of epidemic models at a scale real-world datasets cannot provide.

Comments8 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑