MedTutor: 一种基于检索的LLM系统用于基于病例的医学教育
MedTutor: A Retrieval-Augmented LLM System for Case-Based Medical Education
- Department of Computer Science, Yale University(耶鲁大学计算机科学系)
- Department of Radiology and Biomedical Imaging, Yale School of Medicine(耶鲁医学院放射学与生物医学成像系)
- Interdisciplinary Program for Bioengineering, Seoul National University(首尔国立大学生物工程跨学科项目)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MedTutor是一种基于检索的LLM系统,通过自动从临床病例报告生成教育内容和多项选择题,提升医学教育质量。
AI中文摘要:
医学住院医师的学习过程面临重大挑战,需要既能解读复杂的病例报告,又能快速从可靠来源获取准确的医学知识。住院医师通常会研究病例报告并与同僚和导师讨论,但找到相关教育材料和证据来支持他们从这些病例中学习往往耗时且具有挑战性。为此,我们引入了MedTutor,一种新的系统,旨在通过自动从临床病例报告中生成基于证据的教育内容和多项选择题来增强住院医师的培训。MedTutor利用一种检索增强生成(RAG)流程,以临床病例报告作为输入,生成有针对性的教育材料。该系统的架构特征是混合检索机制,该机制协同查询本地医学教科书和学术文献(使用PubMed、Semantic Scholar API)的最新相关研究,确保生成的内容既基础又当前。检索到的证据通过最先进的重排序模型进行过滤和排序,然后通过LLM生成最终的长文本输出,描述病例报告的主要教育内容。我们对系统进行了严格评估。首先,三位放射科医生评估了输出的质量,发现其具有很高的临床和教育价值。其次,我们通过大规模评估使用LLM作为评判者,以了解LLM能否用于评估系统的输出。我们的分析使用LLM输出与人类专家判断之间的相关性揭示了中等的契合度,并突显了继续需要专家监督的必要性。
英文摘要:
The learning process for medical residents presents significant challenges, demanding both the ability to interpret complex case reports and the rapid acquisition of accurate medical knowledge from reliable sources. Residents typically study case reports and engage in discussions with peers and mentors, but finding relevant educational materials and evidence to support their learning from these cases is often time-consuming and challenging. To address this, we introduce MedTutor, a novel system designed to augment resident training by automatically generating evidence-based educational content and multiple-choice questions from clinical case reports. MedTutor leverages a Retrieval-Augmented Generation (RAG) pipeline that takes clinical case reports as input and produces targeted educational materials. The system's architecture features a hybrid retrieval mechanism that synergistically queries a local knowledge base of medical textbooks and academic literature (using PubMed, Semantic Scholar APIs) for the latest related research, ensuring the generated content is both foundationally sound and current. The retrieved evidence is filtered and ordered using a state-of-the-art reranking model and then an LLM generates the final long-form output describing the main educational content regarding the case-report. We conduct a rigorous evaluation of the system. First, three radiologists assessed the quality of outputs, finding them to be of high clinical and educational value. Second, we perform a large scale evaluation using an LLM-as-a Judge to understand if LLMs can be used to evaluate the output of the system. Our analysis using correlation between LLMs outputs and human expert judgments reveals a moderate alignment and highlights the continued necessity of expert oversight.