发表机构
Scale AI; Emory University; University of California, Santa Cruz; Vanderbilt University Medical Center; Stanford School of Medicine; Stanford Health Care; Clinical Excellence Research Center(Scale AI; 埃默里大学; 加州大学圣克鲁兹分校; 范德堡大学医学中心; 斯坦福医学院; 斯坦福医疗保健系统; 临床卓越研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出首个面向部署的临床智能体基准CliniCARE-Bench,评估智能体对真实纵向电子健康记录的医疗推理,发现原始准确率夸大调查质量,无缺陷准确率更低且会重新排序排行榜。
AI 中文摘要
大型语言模型在医学知识基准测试中表现出色,但可靠的临床部署要求智能体对异构、纵向记录进行可辩护的调查:确定所需证据、检索并整合结构化与自由文本数据、将结论建立在可验证证据之上,以及对无法可靠解决的病例进行弃权(不执行)处理。我们推出CliniCARE-Bench(电子健康记录中医疗推理的临床校准审计),这是一项用于回顾性临床审计的基准:25个经临床医生验证的场景,以基于真实患者的MIMIC-IV数据构建的750个特定患者病例形式呈现。系统通过受管控、可记录的工具环境调查每个病例,该环境用于记录检索、计算和政策访问,并返回四种裁决之一——是、否、不确定:数据缺失,或不确定:医学上模糊——后两者区分了证据缺失与残留的医学模糊性。除裁决准确性外,我们针对由独立多模型裁决产生并经临床委员会审查校准的病例级参考裁决,对患者-证据与政策依据、流程合规性、校准弃权(不执行)、可靠性和效率进行评分。每次检索、计算和报告均可重现,因此调查轨迹可检查和评分。据我们所知,CliniCARE-Bench是首个面向部署的临床智能体基准,在共同的患者级裁决框架内联合评估真实纵向电子健康记录调查、声明级证据依据、监管政策使用、流程合规性及校准弃权(不执行)。在16个智能体系统中,四向准确率介于65.3%至76.1%之间,但原始准确率夸大了调查质量。仅当裁决正确且无禁止的捷径时才给予认可的无缺陷准确率,比原始准确率低4.8至14.8个百分点,并重新排序了排行榜。
英文摘要
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.