arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.03471cs.CLcs.AI

EpiQAL:大型语言模型在流行病学问答与推理中的基准测试

EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning

  • Emory University(埃默里大学)
  • University of Illinois Chicago(伊利诺伊大学香槟分校)
  • Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

Mingyang Wei, Dehai Min, Zewen Liu, Yuzhang Xie, Guanchen Wu, Ziyang Zhang, Carl Yang, Max S. Y. Lau, Qi He, Lu Cheng, Wei Jin

AI总结:

提出EpiQAL基准,通过三个子集(事实回忆、多步推理、不完整信息下结论重建)评估LLM在流行病学推理中的表现,发现当前模型在多步推理上表现有限。

AI中文摘要:

可靠的流行病学推理需要综合研究证据来推断疾病负担、传播动态和人群层面的干预效果。现有的医学问答基准主要强调临床知识或患者层面的推理,但很少有系统评估基于证据的流行病学推理。我们提出了EpiQAL,这是首个针对多种疾病的流行病学问答诊断基准,包含三个从开放获取文献构建的子集。这三个子集逐步测试事实回忆、多步推理以及在不完整信息下的结论重建,并通过结合分类学指导、多模型验证和难度筛选的质量控制流程构建。对涵盖开源和专有系统的15个模型的实验表明,当前LLM在流行病学推理上表现有限,其中多步推理构成最大挑战。模型排名在不同子集间发生变化,仅靠规模并不能预测成功。思维链提示有利于多步推理,但在其他情况下效果不一。EpiQAL为证据基础、推理推理和结论重建提供了细粒度的诊断信号。

英文摘要:

Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clinical knowledge or patient-level reasoning, yet few systematically evaluate evidence-grounded epidemiological inference. We present EpiQAL, to our knowledge the first diagnostic benchmark for epidemiological question answering over research literature, comprising three subsets built from open-access articles across diverse diseases. The three subsets progressively test factual recall, multi-step inference, and conclusion reconstruction under incomplete information, and are constructed through a quality-controlled pipeline combining taxonomy guidance, multi-model verification, and difficulty screening. Experiments on fifteen models spanning open-source and proprietary systems reveal that current LLMs show limited performance on epidemiological reasoning, with multi-step inference posing the greatest challenge. Model rankings shift across subsets, and scale alone does not predict success. Chain-of-Thought prompting benefits multi-step inference but yields mixed results elsewhere. EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction.

补充信息

↑