arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM裁判验证存在,而非不存在:AI临床笔记中的遗漏盲点及恢复方法

LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It

Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris

arXiv 2608.31016首次发表:更新:

AI 中文总结

本研究发现LLM裁判在检测AI临床笔记遗漏时存在盲点,通过重构任务(先列诊疗记录事实再逐一核对)的流水线及GEPA单调用提示可有效恢复检测,且在真实数据上表现优于原有设计。

AI 中文摘要

环境AI书记员起草临床笔记,已发表的审计研究发现其主要错误为遗漏:即诊疗过程中已确认的信息未被记录。标准检查方法是使用LLM裁判:第二个模型对照诊疗记录读取笔记并标记问题。本研究探究裁判是否能检测遗漏,因公开语料库的临床医生参考笔记与诊疗记录存在实质性差异,无法提供答案密钥。本基准包含500条来自审计事实表的单错误笔记对,其中298条存在明确遗漏的命名事实,202条为添加或修改的对照内容。在8种裁判设计中,配对判别(缺陷笔记评分低于其完整对照,0.5为随机水平)在添加或修改内容上的表现为0.79-0.94,在遗漏内容上仅为0.50-0.63;对单条笔记而言,无设计能可靠标记遗漏的频率高于完美笔记。措辞修改、投票及GEPA提示优化可调整操作点,但无法形成可用检测。重构任务可恢复检测效果:列出诊疗记录确认的所有事实,再对照笔记检查每条事实。两种方法独立实现该效果且各有优劣:按事实划分的流水线,及GEPA演化的单调用提示。流水线标记缺失事实及其严重程度,误报率为2.7%;单调用检测率更高(36.9% vs 24.6%,p=0.002),误报率为6.2%,每条笔记成本仅为流水线的十分之一。一名医师作者验证了70个项目,在两种方法分歧的10个案例中均支持流水线(p=0.002);另一名非作者临床医生盲评严重程度 rubric,结果与流水线的一致性在一个等级范围内。在配套普查的真实供应商笔记中,无基准阈值可迁移,但重新校准后的单调用检测率高于8种设计中最优者,误报率仅为其一半。若遗漏的事实在其他地方被重述,则两种方法均无法检测。本研究发布了该基准、提示及判断结果。

英文摘要

Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the note fails to record. The standard check is an LLM judge: a second model reads the note against the transcript and flags problems. We ask whether judges detect omissions. Public corpora cannot supply the answer key: their clinician reference notes and transcripts are materially discrepant. Our benchmark has 500 single-error note pairs from audited fact sheets, 298 with a named fact certainly absent and 202 added-or-altered controls. Across eight judge designs, paired discrimination (the flawed note below its clean twin, 0.5 a coin flip) reads 0.79-0.94 on added or altered content and 0.50-0.63 on omissions. On single notes, no design flags omissions reliably more often than perfect notes. Wording changes, voting and GEPA prompt optimisation move the operating point without creating usable detection. Restructuring the task recovers it: list the facts the transcript establishes, then check the note for each. Two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call. The pipeline's flags name the missing fact and its severity at 2.7% false alarms. The single call detects more (36.9% against 24.6%, p=0.002) at 6.2% false alarms and a tenth of the cost per note. A physician author validated 70 items and, where the two routes disagree, sided with the pipeline on 10 of 10 (p=0.002). A second clinician, not an author, graded the severity rubric blind and agrees to within a grade. On real vendor notes from a companion census no benchmark threshold transfers, but the re-calibrated single call detects more than the best of the eight at half its false-alarm rate. Omissions whose fact is restated elsewhere defeat both routes. We release the benchmark, prompts and judgements.

Comments97 pages, 8 figures. Dataset: https://huggingface.co/datasets/ComposoAI/OmissionBench . Code: https://github.com/composo-ai/omission-bench . Companion paper: "One note in three: a verified census of three deployed AI scribes, and the instrument that counted it"

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑