评估并增强大语言模型在罕见病问答中的表现
Assessing and Enhancing Large Language Models in Rare Disease Question-answering
- Rice University(莱斯大学)
- Baylor College of Medicine(贝勒医学院)
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究构建ReDis-QA罕见病问答数据集评估大语言模型的罕见病诊断能力,同时搭建首个罕见病语料库ReCOP,可平均提升模型准确率8%并增强答案可信度与可追溯性。
AI中文摘要:
尽管大语言模型(LLMs)在普通医学领域展现出出色能力,但它们在罕见病诊断方面的表现仍存疑问。为解答这一问题,我们旨在评估LLMs在罕见病中的诊断性能,并探索提升其在该领域有效性的方法。本研究中,我们构建了罕见病问答数据集ReDis-QA,用于评估LLMs的罕见病诊断性能。具体而言,我们在ReDis-QA数据集中收集了1360个高质量问答对,覆盖205种罕见病。此外,我们为每个问题标注了元数据,便于提取任意特定疾病及其属性对应的子集。基于ReDis-QA数据集,我们对多款开源LLMs进行基准测试,结果显示诊断罕见病对这些模型而言仍是重大挑战。\n为助力罕见病诊断的检索增强生成,我们从美国国家罕见疾病组织(NORD)数据库中采集了首个罕见病语料库ReCOP。具体来说,我们将每种罕见病的报告拆分为多个文本块,每个块对应疾病的不同属性,包括概述、症状、病因、影响、相关疾病、诊断以及标准疗法。这种结构确保每个文本块内的信息与问题保持一致匹配。实验结果表明,ReCOP可将LLMs在ReDis-QA数据集上的准确率平均提升8%。此外,它还能显著引导LLMs生成可追溯至现有文献的可信答案与解释。
英文摘要:
Despite the impressive capabilities of Large Language Models (LLMs) in general medical domains, questions remain about their performance in diagnosing rare diseases. To answer this question, we aim to assess the diagnostic performance of LLMs in rare diseases, and explore methods to enhance their effectiveness in this area. In this work, we introduce a rare disease question-answering (ReDis-QA) dataset to evaluate the performance of LLMs in diagnosing rare diseases. Specifically, we collected 1360 high-quality question-answer pairs within the ReDis-QA dataset, covering 205 rare diseases. Additionally, we annotated meta-data for each question, facilitating the extraction of subsets specific to any given disease and its property. Based on the ReDis-QA dataset, we benchmarked several open-source LLMs, revealing that diagnosing rare diseases remains a significant challenge for these models. To facilitate retrieval augmentation generation for rare disease diagnosis, we collect the first rare diseases corpus (ReCOP), sourced from the National Organization for Rare Disorders (NORD) database. Specifically, we split the report of each rare disease into multiple chunks, each representing a different property of the disease, including their overview, symptoms, causes, effects, related disorders, diagnosis, and standard therapies. This structure ensures that the information within each chunk aligns consistently with a question. Experiment results demonstrate that ReCOP can effectively improve the accuracy of LLMs on the ReDis-QA dataset by an average of 8%. Moreover, it significantly guides LLMs to generate trustworthy answers and explanations that can be traced back to existing literature.