arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型可提高医师准确率,但会导致错误依赖

Large language models improve physician accuracy but lead to false reliance

Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker

arXiv 2608.00817首次发表:更新:

AI 中文总结

本研究开发智能体式检索增强型LLM CORA,发现其可将46名医师的诊断准确率从70.8%提升至82.6%,但存在错误答案获引用支持时医师抵抗大幅下降的安全风险。

AI 中文摘要

检索增强型大型语言模型(LLMs)可提供关联来源的临床支持,但其价值取决于展示的证据是引导还是扭曲医师的依赖。我们开发了CORA,一种智能体式检索增强型LLM,以研究关联来源的辅助如何影响医师决策。CORA保持了基准性能,且在模型训练数据截止后发布的病例上取得了更大提升。在一项包含46名医师的研究中,无辅助时准确率为70.8%,使用CORA后升至82.6%。支持性引用可预测正确答案(87.7%对65.5%),但引用造成了重要的不对称性:感知到的支持使医师采纳正确建议的比例从34%升至76.9%,但当LLM的错误答案有引用支持时,医师对其的抵抗从92%降至34.8%。这些发现表明,关联来源的LLM辅助可提高医师准确率,同时引入了依赖于依据的安全风险。

英文摘要

Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑