arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EpiWorld:将LLM策略智能体锚定于流行病学世界模型

EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models

Zeeshan Memon, Yiqi Su, Kai Shu, Naren Ramakrishnan, Liang Zhao

arXiv 2610.02744首次发表:更新:

发表机构

Emory University; Virginia Tech(埃默里大学; 弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EpiWorld通过将LLM策略代理与流行病学世界模型及分层技能库结合,实现干预后果预测与政策优化,在回顾性数据上显著降低住院率,优于现有基线。

AI 中文摘要

流行病干预政策是文本制品,人类决策者通过自然语言对其进行解读、论证和修订,这使得大型语言模型成为流行病政策推理的自然候选者。然而,一个朴素的LLM缺乏预测干预后果所需的流行病动力学、评估严重程度所需的定量监测信号,以及界定可允许行为的制度约束。我们提出EpiWorld,一个闭环框架,将LLM策略执行者锚定于一个学习到的动作条件流行病学世界模型和一个分层技能库,该技能库包含公共卫生协议、监测工具以及通过事后分析积累的适应性经验教训。给定一个候选干预措施,世界模型预测区域流行病演变,并实现快速反事实推演,为政策选择和优化提供反馈。模拟未来的结果被提炼为可复用的经验教训,同时协议约束保持不变,使得决策过程在不牺牲可解释性或可控性的前提下得以改进。我们在回顾性COVID-19和流感数据集上评估了世界模型和端到端框架:世界模型在所有预测基线中取得了最佳分布外峰值平均绝对误差(Peak-MAE),闭环框架在数据集上将累计住院人数最多减少59%,在六个LLM骨干网络上平均减少约16%,优于强化学习和最优控制策略基线。

英文摘要

Epidemic intervention policies are textual artefacts that human decision-makers interpret, justify, and revise through natural language, making large language models a natural candidate for epidemic policy reasoning. A naive LLM, however, lacks the epidemic dynamics needed to project intervention consequences, the quantitative surveillance signals required to assess severity, and the institutional constraints that define admissible actions. We present EpiWorld, a closed-loop framework that grounds an LLM policy actor in a learned action-conditioned epidemiological world model and a tiered skill library of public-health protocols, surveillance tools, and adaptive lessons accumulated through after-action analysis. Given a candidate intervention, the world model predicts regional epidemic evolution and enables fast counterfactual rollouts that provide feedback for policy selection and refinement. Outcomes of simulated futures are distilled into reusable lessons while protocol constraints remain fixed, allowing the decision process to improve without sacrificing interpretability or controllability. We evaluate both the world model and the end-to-end framework on retrospective COVID-19 and Influenza datasets: the world model achieves the best out-of-distribution Peak-MAE among all forecasting baselines, and the closed-loop framework reduces cumulative hospitalisation by up to 59% across datasets and by an average of ~16% across six LLM backbones, outperforming reinforcement-learning and optimal-control policy baselines.

CommentsAccepted to Findings of EMNLP 2026. 22 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑