arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

感知与敏感性:基于医生专家基准评估大语言模型临床分诊建议

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi

arXiv 2609.38600首次发表:更新:

发表机构

Massachusetts Institute of Technology; Worcester Polytechnic Institute(麻省理工学院; 伍斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建含6,000余临床场景的基准,发现大语言模型在临床分诊中比医生更易推荐不必要护理,且对性别和语气扰动更敏感,凸显部署前需以医生行为为基准进行评估。

AI 中文摘要

随着大语言模型(LLMs)在临床环境中日益广泛的应用,评估其在临床文本真实变化下的可靠性变得至关重要。我们在临床分诊场景中研究这一问题,在保持潜在临床情境不变的前提下,通过文本扰动将大语言模型与执业医师进行比较。我们构建了一个包含超过6,000个临床场景、7,000条医生标注和225,000条模型响应的基准测试。基于该基准,我们得出两个关键发现。首先,在基线条件下,大语言模型比医生更倾向于推荐不必要的护理,且这种倾向在扰动输入下进一步增强。此外,我们发现大语言模型的推荐对性别和语气扰动的敏感性高于人类推荐。综合来看,这些结果表明大语言模型可能因临床无关的文本变化而产生不同输出,凸显了以执业医师行为为基础、面向实际部署的评估的必要性。

英文摘要

As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting. We introduce a benchmark of over 6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses. Using this benchmark, we make two key observations. First, LLMs are more likely than physicians to recommend unnecessary care at baseline, and this tendency increases under perturbed inputs. Further, we find that LLM recommendations are more sensitive to gender and tone perturbations than human recommendations. Together, these results demonstrate that LLMs can vary under clinically irrelevant textual changes, highlighting the need for deployment-oriented evaluations grounded in expert physician behavior.

CommentsAccepted to Findings EMNLP 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑