发表机构
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出医学因果假设验证的LLM评估框架,测试8个LLM在17个医学因果假设上的性能,发现其虽召回率高,但提供有效证据及拒绝无依据假设的表现差,凸显需严格评估LLM在医疗领域的应用。
AI 中文摘要
大型语言模型(LLMs)在搜索和信息检索中的应用日益广泛,这凸显了评估其在医疗等高风险领域可靠性的必要性。尽管LLMs能有效回答疾病、症状和治疗相关问题,但其准确评估因果关系并将结论建立在已验证科学证据上的能力仍不明确。本文开展一项初步小规模研究,调查LLMs在评估医学因果主张及用同行评审研究支撑主张的准确性。我们提出一个用于因果假设验证的评估框架,可系统跟踪现有及未来LLMs的性能。我们评估8个LLMs在17个医学因果假设上的性能,以检验其能否利用文献中的科学证据可靠验证这些假设。我们依据6项标准对其提供的科学证据进行系统标注(共1067个标注点),并采用9项评估指标进行评估。分析显示,LLMs虽表现出较强的召回率,但在提供有效科学文章及支撑证据、拒绝无依据假设方面常表现不佳。这些发现凸显当前LLMs的关键局限,即其尚不能完全信赖以验证生物医学文献中的因果关系。本研究强调,在医疗场景中使用LLMs进行搜索和检索前,需进行严格评估。
英文摘要
The growing use of large language models (LLMs) for search and information retrieval underscores the need to evaluate their reliability in high-stakes domains such as healthcare. Although LLMs can effectively answer questions about diseases, symptoms, and treatments, their ability to accurately assess causal relationships and ground their conclusions in verified scientific evidence remains unclear. Here, we present a preliminary, small-scale study that investigates the accuracy of LLMs in evaluating causal medical claims and supporting them with peer-reviewed research. We propose an evaluation framework for causal hypothesis verification that can be used to systematically track the performance of existing and future LLMs. We assess the performance of eight LLMs on 17 medical causal hypotheses to evaluate whether they can reliably verify these hypotheses using scientific evidence from the literature. We systematically annotate the scientific evidence they provide according to six criteria (a total of 1,067 annotation points) and assess them with nine evaluation metrics. Our analysis shows that while LLMs exhibit strong recall, they often perform poorly at providing valid scientific articles and evidence for support and at rejecting unsupported hypotheses. These findings highlight a critical limitation of current LLMs, as they cannot yet be trusted fully to verify causal relationships from the biomedical literature. This work underscores the need for rigorous evaluation before using LLMs for search and retrieval in healthcare settings.
Journal refCONSEQUENCES Workshop @ RecSys '26, October 02, 2026, Minneapolis, MN, USA