arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向更好评估大型语言模型(LLMs)在临床错误检测中的性能

Toward Better Assessment of LLMs' Performance in Clinical Error Detection

Yifan Zhang, Rahmatollah Beheshti

arXiv 2608.16643首次发表:更新:

发表机构

University of Delaware(特拉华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLMs临床错误检测评估的不足,通过多语言多模型实验发现多数LLMs配对判别能力差,提出用配对评估补充聚合指标的改进方案。

AI 中文摘要

临床文档错误的自动检测是大型语言模型(LLMs)的一个有前景的应用,然而部署此类模型的决策依赖于对每份临床笔记单独评估的基准。错误检测基准通常通过向笔记中注入错误来构建,使得每份错误笔记都有一个对应的自然版本。聚合判别指标(如平衡准确率或F1)并未利用这种结构,我们表明这种遗漏会产生严重后果。具体而言,我们在3种语言的4个标准化临床错误检测测试集上评估了15种不同的LLMs,发现15个模型中有13个低于随机配对判别水平,尽管它们达到了标准实践会视为中等的F1分数。我们还观察到,潜在的偏差模式因语言而异:同一模型可能在一种语言上默认“无错误”,而在另一种语言上过度标记错误。为了诊断判别失效的位置,我们进一步引入了一种对模型输出中引用的证据进行评分的方法。我们发现,虽然模型始终能定位与错误相关的内容,但它们无法对干净的对应内容给出相应的正确判决。最后,我们表明F1和配对准确率受相同潜在偏差的驱动方向相反,因此按F1对模型排序可能会系统性地提升最弱的判别器。对于安全关键的临床自然语言处理(NLP)应用,我们主张在基准报告中用配对评估补充聚合指标。代码和分析脚本可在此https URL获取。

英文摘要

Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart. Aggregate discriminative metrics (e.g., balanced accuracy or F1) do not exploit this structure. We show that this omission is consequential. In particular, evaluating 15 diverse LLMs on 4 standardized clinical error-detection test sets across 3 languages, we find that 13 of 15 models fall below the level of random pairwise discrimination, even while achieving F1 scores that standard practice would read as moderate. We also observe that the underlying bias patterns differ across languages: the same model can default to "no error" on one language and over-flag errors on another. To diagnose where discrimination breaks down, we further introduce a procedure to score the evidence models cite in their outputs. We find that while models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart. Finally, we show that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators. For safety-critical clinical NLP applications, we advocate for supplementing aggregate metrics with paired evaluations in benchmark reporting. Code and analysis scripts are available at https://github.com/healthylaife/paired-clinical-eval.

CommentsAccepted at Machine Learning for Healthcare (MLHC) 2026; to appear in Proceedings of Machine Learning Research (PMLR), Vol. 340

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑