arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2309.09362cs.CL

语言模型在医疗应用中易受患者错误自我诊断的影响

Language models are susceptible to incorrect patient self-diagnosis in medical applications

  • University of Maryland, College Park(马里兰大学帕克分校)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Rojin Ziaei, Samuel Schmidgall

更新

AI总结:

本文通过将美国医学委员会考试题修改为包含患者自我诊断报告,评估多种LLMs,发现患者提供错误偏倚验证信息会显著降低LLMs的诊断准确性,揭示其对自我诊断错误的高度易感性。

AI中文摘要:

大型语言模型(LLMs)作为医疗保健的潜在工具正变得越来越重要,它们有助于临床医生、研究人员和患者之间的沟通。然而,传统上对LLMs在医学考试题目上的评估并不能反映真实医患互动的复杂性。这种复杂性的一个例子是患者自我诊断的引入,即患者试图从各种来源诊断自己的医疗状况。虽然患者有时会得出准确的结论,但由于患者过度强调偏倚验证信息,他们更常被引向误诊。在这项工作中,我们向多种LLMs提供了来自美国医学委员会考试的多项选择题,这些题目经过修改,加入了患者的自我诊断报告。我们的研究结果强调,当患者提出错误的偏倚验证信息时,LLMs的诊断准确性会急剧下降,揭示了其在自我诊断中对错误的高度易感性。

英文摘要:

Large language models (LLMs) are becoming increasingly relevant as a potential tool for healthcare, aiding communication between clinicians, researchers, and patients. However, traditional evaluations of LLMs on medical exam questions do not reflect the complexity of real patient-doctor interactions. An example of this complexity is the introduction of patient self-diagnosis, where a patient attempts to diagnose their own medical conditions from various sources. While the patient sometimes arrives at an accurate conclusion, they more often are led toward misdiagnosis due to the patient's over-emphasis on bias validating information. In this work we present a variety of LLMs with multiple-choice questions from United States medical board exams which are modified to include self-diagnostic reports from patients. Our findings highlight that when a patient proposes incorrect bias-validating information, the diagnostic accuracy of LLMs drop dramatically, revealing a high susceptibility to errors in self-diagnosis.

补充信息

↑