检索增强医疗问答中检索流程设计的系统研究
A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering
- Department of Mechatronics & Industrial Engineering, Chittagong University of Engineering & Technology(吉大港工程与技术大学机电与工业工程系)
- Department of Mechanical Engineering, Chittagong University of Engineering & Technology(吉大港工程与技术大学机械工程系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文系统评估了检索增强医疗问答的性能,发现检索增强显著提升了零样本医疗问答效果,最佳配置为密集检索加查询改写和重排序,准确率达60.49%。
AI中文摘要:
大型语言模型(LLMs)在医疗问答中表现出强大能力;然而,纯参数模型常面临知识缺口和事实基础不足的问题。检索增强生成(RAG)通过将外部知识检索整合到推理过程中解决这一限制。尽管RAG医疗系统日益受到关注,但个体检索组件对性能的影响仍不明确。本文通过MedQA USMLE基准和结构化教材知识库,系统评估了检索增强医疗问答。分析了语言模型、嵌入模型、检索策略、查询改写和交叉编码器重排序之间的交互,包含四十种配置的统一实验框架。结果表明,检索增强显著提升了零样本医疗问答性能。最佳配置为密集检索加查询改写和重排序,准确率达60.49%。领域专用语言模型比通用模型更有效利用检索到的医疗证据。分析还揭示了检索效果与计算成本之间的明确权衡,简单密集检索配置在保持高吞吐量的同时提供强性能。所有实验均在单块消费级GPU上完成,证明在有限计算资源下可以系统评估检索增强医疗问答系统。
英文摘要:
Large language models (LLMs) have demonstrated strong capabilities in medical question answering; however, purely parametric models often suffer from knowledge gaps and limited factual grounding. Retrieval-augmented generation (RAG) addresses this limitation by integrating external knowledge retrieval into the reasoning process. Despite increasing interest in RAG-based medical systems, the impact of individual retrieval components on performance remains insufficiently understood. This study presents a systematic evaluation of retrieval-augmented medical question answering using the MedQA USMLE benchmark and a structured textbook-based knowledge corpus. We analyze the interaction between language models, embedding models, retrieval strategies, query reformulation, and cross-encoder reranking within a unified experimental framework comprising forty configurations. Results show that retrieval augmentation significantly improves zero-shot medical question answering performance. The best-performing configuration was dense retrieval with query reformulation and reranking achieved 60.49% accuracy. Domain-specialized language models were also found to better utilize retrieved medical evidence than general-purpose models. The analysis further reveals a clear tradeoff between retrieval effectiveness and computational cost, with simpler dense retrieval configurations providing strong performance while maintaining higher throughput. All experiments were conducted on a single consumer-grade GPU, demonstrating that systematic evaluation of retrieval-augmented medical QA systems can be performed under modest computational resources.