AI 中文总结
本研究以MMTV-LV与乳腺癌为案例,构建数据集测试4款LLM在24篇微生物致癌研究论文的证据提取与评估任务上的表现,发现GPT-5等两款LLM性能与专家相当,可用于自动化系统证据综合,但仍存在方法学评估等弱点。
AI 中文摘要
已确认的致癌微生物对癌症负担有显著贡献,识别新型微生物致癌性可产生降低疾病负担的策略,但相关证据分散,人类难以全面综合。大型语言模型(LLM)或可实现规模化、专家级的系统证据综合以识别微生物-癌症配对,但此类能力尚未得到验证。本研究招募领域专家创建数据集,以Gemini 2.5 Pro、Gemini 2.5 Flash、GPT-5、GPT-5 Nano这4款LLM为对象,以MMTV-LV与乳腺癌为案例,针对24篇研究论文开展性能基准测试。我们设计了一套包含多项选择题、李克特量表、多选及自由文本题型的结构化证据提取与评估模板(24篇论文共77个条目),采用新型指标针对每个问题实例确定(1)专家间的一致性,以及(2)专家与各LLM间的一致性。通过比较专家间、专家-LLM的一致性分布,评估LLM是否可作为额外专家提升或维持专家间一致性,还对自由文本回答进行了定性评估。在所有题型中,LLM回答与专家高度一致,其中GPT-5与GPT-5 Nano的得分分布与专家无差异;Gemini模型表现相似,但在应用微生物致癌性标准时明显更宽松,幻觉情况罕见。方法学评估及全文矛盾识别是LLM最持久的漏洞,GPT-5与GPT-5 Nano在结构化领域研究论文评估任务上与专家无差异,这支持将LLM用于自动化系统证据综合,但方法学评估任务与全文矛盾识别仍是需加强的弱点。
英文摘要
Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensively synthesize. LLMs may enable scalable, expert-level systematic evidence synthesis to identify microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Domain experts were recruited to create a dataset to benchmark LLM performance (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano) on 24 research papers using MMTV-LV and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal, consisting of MCQ, Likert-scale, multi-select, and free-text question types (77 items across 24 papers). Agreement between (1) experts and (2) experts and each LLM was determined per question instance using novel metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement distributions to determine whether LLMs behaved as additional experts by increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Across all question types, LLM responses aligned closely with experts, with GPT-5 and GPT-5 Nano achieving score distributions indistinguishable from experts. Gemini models behaved similarly but were significantly more lenient in applying microbial oncogenesis criteria. Hallucinations were rare. Methodological appraisal and identification of contradictions within full-texts were the most persistent LLM vulnerabilities. GPT-5 and GPT-5 Nano were indistinguishable from experts on structured domain research paper evaluation tasks. This supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-texts remain weaknesses requiring strengthening.
CommentsPublished in Frontiers in Cellular and Infection Microbiology, 45 pages, 14 figures
Journal refKokkas K, Wang H, Klein R, et al. (2026) Artificial intelligence can match domain experts in evidence extraction and critical appraisal of microbial oncogenesis research publications. Front. Cell. Infect. Microbiol. 16:1876326
DOI:10.3389/fcimb.2026.1876326