arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型临床推理评估的评分标准全景:现状、缺失与需整合之处

A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

arXiv 2610.01938首次发表:更新:

发表机构

DRIVE-Health CDT; Department of Biostatistics and Health Informatics; Institute of Psychiatry, Psychology and Neuroscience; King’s College London; Cleveland Clinic London(DRIVE-Health 博士训练中心; 生物统计学与健康信息学系; 精神病学、心理学与神经科学研究所; 伦敦国王学院; 伦敦克利夫兰诊所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统梳理医学教育与临床LLM评估文献,提出覆盖六维度的临床推理评估框架,指出需整合现有工具并强化忠实性等薄弱环节。

AI 中文摘要

考试式的准确性并不能证明大型语言模型(LLM)在临床记录上的推理能力是否良好。我们将临床推理定义为跨时间和来源整合并更新证据,以形成、修正和论证患者的问题表征及可辩护的计划。本结构化叙述性综述梳理了三类文献:医学教育评估工具、2023年起发表的临床LLM基准,以及评估长文本生成的通用领域方法。我们考察了六个维度:问题表征、时间综合、鉴别与管理推理、反事实推理、校准的不确定性,以及推理的忠实性。预印本被纳入并加以标注。没有任何单一工具覆盖全部六个维度。问题表征以及鉴别或管理推理得到了较为充分的覆盖,尽管可靠性因工具和场景而异。TIMER-Eval针对时间综合,ER-Reason评估序贯诊断信念更新。专门的不确定性和反事实评估正在兴起,但其对纵向自由文本推理的适用性仍然有限。事实完整性在通用领域评估中理论化程度较高,并有早期临床证据表明存在重要遗漏。忠实性仍是最薄弱的维度,仅有一项针对多项选择题的临床因果消融研究被识别。现有工具应通过二元评分条目、独立的完整性与正确性评分、具有不可补偿安全上限的病例特定重要性加权、时间顺序一致性检查,以及机会校正的可靠性报告来加以整合。针对纵向自由文本记录的校准不确定性、反事实推理和忠实性,仍需进一步的设计工作。本综述提供了设计依据,而非经过验证的工具。

英文摘要

Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.

Comments13 pages, 1 table. Structured narrative review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑