预测正确,推理错误:揭示大语言模型在类风湿关节炎疾病诊断中的错位
Right Prediction, Wrong Reasoning: Uncovering LLM Misalignment in RA Disease Diagnosis
- KIIT Bhubaneswar(基特布巴内斯瓦尔大学)
- KIMS Bhubaneswar(布巴内斯瓦尔 KIMS 医院)
- Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究大语言模型在类风湿关节炎诊断中的表现,发现其预测准确率约95%但近68%的推理错误,揭示了预测与解释之间的严重错位,警示临床中不能依赖LLM的解释。
AI中文摘要:
大语言模型(LLMs)提供了一种有前景的预筛查工具,可改善早期疾病检测,并为医疗服务不足的社区提供更好的医疗保健可及性。各种疾病的早期诊断在医疗保健中仍然是一项重大挑战,主要原因是早期症状的非特异性、专业医疗从业人员的短缺以及需要长时间的临床评估,这些都可能延误治疗并对患者预后产生不利影响。凭借在一系列疾病预测中令人印象深刻的准确性,LLMs 有潜力彻底改变各种医疗状况的临床预筛查和决策。在这项工作中,我们使用真实世界患者数据研究 LLMs 对类风湿关节炎(RA)的诊断能力。患者数据与医学专家的诊断一同收集,并将 LLMs 的性能与专家对 RA 疾病预测的诊断进行比较评估。我们注意到疾病诊断中一个有趣的模式,并发现预测与解释之间存在意外的错位。我们使用不同的 LLM 智能体进行了一系列多轮分析。表现最佳的模型预测类风湿关节炎(RA)疾病的准确率约为 95%。然而,当医学专家评估该模型生成的推理时,他们发现近 68% 的推理是不正确的。这项研究凸显了 LLMs 高预测准确率与其有缺陷推理之间的明显错位,引发了关于在临床环境中依赖 LLM 解释的重要问题。LLMs 提供了不正确的推理却得出了 RA 疾病诊断的正确答案。
英文摘要:
Large language models (LLMs) offer a promising pre-screening tool, improving early disease detection and providing enhanced healthcare access for underprivileged communities. The early diagnosis of various diseases continues to be a significant challenge in healthcare, primarily due to the nonspecific nature of early symptoms, the shortage of expert medical practitioners, and the need for prolonged clinical evaluations, all of which can delay treatment and adversely affect patient outcomes. With impressive accuracy in prediction across a range of diseases, LLMs have the potential to revolutionize clinical pre-screening and decision-making for various medical conditions. In this work, we study the diagnostic capability of LLMs for Rheumatoid Arthritis (RA) with real world patients data. Patient data was collected alongside diagnoses from medical experts, and the performance of LLMs was evaluated in comparison to expert diagnoses for RA disease prediction. We notice an interesting pattern in disease diagnosis and find an unexpected \textit{misalignment between prediction and explanation}. We conduct a series of multi-round analyses using different LLM agents. The best-performing model accurately predicts rheumatoid arthritis (RA) diseases approximately 95\% of the time. However, when medical experts evaluated the reasoning generated by the model, they found that nearly 68\% of the reasoning was incorrect. This study highlights a clear misalignment between LLMs high prediction accuracy and its flawed reasoning, raising important questions about relying on LLM explanations in clinical settings. \textbf{LLMs provide incorrect reasoning to arrive at the correct answer for RA disease diagnosis.}