三分一的笔记:对三款已部署AI记录员的验证普查及计数工具
One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
浏览论文内容
中文总结 AI 辅助
本研究审计三款商用AI记录员在142次问诊中的表现,发现31.3%的笔记存在经验证的故障,故障率受记录员和计数工具共同影响,公开了相关发现、证据及可复现流程。
中文摘要 AI 辅助
环境AI记录员在临床医生会对每一份笔记签字确认的前提下起草临床笔记。我们对三款商用AI记录员在相同的142次问诊中进行了审计,这些问诊包括565份来自英国初级保健和美国门诊就诊的记录,以及虚构场景。十二次探索性审核提出了13678个候选错误;通过重要性筛选的5898个错误被提交给由两个不同系列模型组成的对抗性小组,每个模型都被要求反驳其能发现的错误,最终有618个错误留存。三分一的笔记(31.3% [27.0, 35.6])存在经验证的故障,集中在过敏和药物信息、虚构的患者身份,以及电话问诊中本不存在的检查内容却被记录为病史。未向任何产品提供患者记录;排除会预先填充记录的两类情况(虚构身份和日期)后,故障发生率为24.8% [20.8, 29.0]。有一种故障模式不符合我们的方案,它来自已发表的记录员错误分类:临床医生撤回的治疗被记录为已实施的护理。两名临床医生对盲法、不重叠的样本进行了裁决:作为作者的医生支持21项发现中的20项(95.2% [77.3, 99.2]),而非作者的独立临床医生支持12项中的12项([75.8, 100]);两人均判定所有抽样的拒绝均为真实。故障发生率既取决于记录员,也取决于计数工具。在模型、证据和设置固定的情况下,仅审核指令就可使候选错误的验证比例从9.3%提升至79.0%,而审核模型系列也会影响该比例:在该审核指令下,温和模型单独标记54.8%的笔记,而另一模型标记27.8%。根据标准不同,抽样笔记中存在故障的比例在28%至97%之间。已发表的审计之间存在分歧,其差异幅度仅由工具差异即可产生:遗漏错误占已发表审计错误的54%-86%,而我们的遗漏错误占比为23.1%。我们公开了全部618项发现及转录文本旁的证据、所有提示和模型版本,以及可复现的流程。
英文摘要
Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.