arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2405.11407cs.CLcs.LG

公共大语言模型能否用于医疗状况的自我诊断?

Can Public LLMs be used for Self-Diagnosis of Medical Conditions ?

  • The University of Texas at Tyler(德克萨斯大学泰勒分校)

机构由 AI 辅助整理,请以论文原文为准。

Nikil Sharan Prabahar Balasubramanian, Sagnik Dakshit

更新

AI总结:

本文通过构建10000个样本的提示工程数据集,评估公共大语言模型GPT-4.0和Gemini在医疗自我诊断任务中的表现,发现二者准确率分别为63.07%和6.01%,并探讨了其挑战、局限及检索增强生成带来的性能提升。

AI中文摘要:

深度学习的进步引发了人们对基础深度学习模型开发的大规模兴趣。大语言模型(LLM)的发展已演变为对话任务中的一种变革性范式,这促使其在医疗保健这一关键领域中也得到整合与扩展。随着LLM日益普及,并通过开源模型以及与其他应用的集成实现公共访问,有必要研究其潜力与局限性。其中一项关键任务——基于症状进行医疗状况的自我诊断——关系到公共利益,LLM虽已被应用于该任务,但仍需更深入的理解。Gemini与谷歌搜索、GPT-4.0与必应搜索的广泛整合,已使自我诊断的趋势从使用搜索引擎转向使用对话式LLM模型。鉴于该任务的关键性质,审慎地调查并理解公共LLM在自我诊断任务中的潜力与局限性是明智之举。在本研究中,我们构建了一个包含10000个样本的提示工程数据集,并在自我诊断这一通用任务上测试了性能。我们比较了当前最先进的GPT-4.0和免费的Gemini模型在自我诊断任务上的表现,记录到的准确率分别为63.07%和6.01%,形成鲜明对比。我们还讨论了Gemini和GPT-4.0在自我诊断任务中所面临的挑战、局限性以及潜力,以促进未来研究,并推动其对公众知识产生更广泛的影响。此外,我们展示了使用检索增强生成(Retrieval Augmented Generation)在自我诊断任务上的潜力与性能提升。

英文摘要:

Advancements in deep learning have generated a large-scale interest in the development of foundational deep learning models. The development of Large Language Models (LLM) has evolved as a transformative paradigm in conversational tasks, which has led to its integration and extension even in the critical domain of healthcare. With LLMs becoming widely popular and their public access through open-source models and integration with other applications, there is a need to investigate their potential and limitations. One such crucial task where LLMs are applied but require a deeper understanding is that of self-diagnosis of medical conditions based on bias-validating symptoms in the interest of public health. The widespread integration of Gemini with Google search and GPT-4.0 with Bing search has led to a shift in the trend of self-diagnosis using search engines to conversational LLM models. Owing to the critical nature of the task, it is prudent to investigate and understand the potential and limitations of public LLMs in the task of self-diagnosis. In this study, we prepare a prompt engineered dataset of 10000 samples and test the performance on the general task of self-diagnosis. We compared the performance of both the state-of-the-art GPT-4.0 and the fee Gemini model on the task of self-diagnosis and recorded contrasting accuracies of 63.07% and 6.01%, respectively. We also discuss the challenges, limitations, and potential of both Gemini and GPT-4.0 for the task of self-diagnosis to facilitate future research and towards the broader impact of general public knowledge. Furthermore, we demonstrate the potential and improvement in performance for the task of self-diagnosis using Retrieval Augmented Generation.

补充信息

↑