arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用大型语言模型开展疾病传播模型的系统文献综述

Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models

Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak

arXiv 2608.26150首次发表:更新:

发表机构

George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发了一套用于从536篇基于智能体的建模论文中提取模型相关信息的LLMs pipeline,对比了GPT-4.1、GPT-5.0与人工开展的SLR结果,为prompt开发提供了见解,凸显了LLMs用于全规模SLR的潜力与局限性。

AI 中文摘要

大型语言模型(LLMs)的最新进展为简化乃至自动化诸多研究流程(包括系统文献综述SLR)创造了新机遇。本研究报告了一套LLMs pipeline的开发工作,用于从536篇同行评审的基于智能体的建模论文中提取与模型相关的信息,并将结果与人工开展的SLR结果进行对比。结果显示,GPT-4.1的论文级准确率约为77.95%,GPT-5.0的论文级准确率约为81.67%;领域级准确率范围为32.40%至100.00%,其中更复杂或主观的领域表现可靠性较低。重要的是,研究发现LLMs之间的一致性可作为输出质量的潜在指标:低一致性可能意味着存在幻觉,而高一致性结合低准确率则可能指向人工数据集存在噪声或错误。总体而言,本研究为prompt开发提供了实用见解,并凸显了在建模与仿真领域使用LLMs开展全规模SLR的潜力与局限性。

英文摘要

Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extracting model-relevant information from 536 peer-reviewed agent-based modeling papers. We compare the results with those of a human-conducted SLR. Our results show paper-level accuracies of approximately 77.95% for GPT-4.1 and 81.67% for GPT-5.0. Field-level accuracy ranges from 32.40% to 100.00%, with more complex or subjective fields performing less reliably. Importantly, we find that agreement between LLMs is a potential indicator of output quality: low agreement may signal hallucinations, whereas high agreement combined with low accuracy may point to noise or errors in the human dataset. Overall, our study provides practical insights into prompt development and highlights both the potential and limitations of using LLMs for full-scale SLRs in the modeling and simulation domain.

CommentsTo be published in the Winter Simulation Conference 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑