arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于评估大语言模型回答中表达性临床推理的评分量规提案

A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

arXiv 2609.37788首次发表:更新:

发表机构

Cleveland Clinic London(克利夫兰诊所伦敦分院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种多维评分量规,融合医学教育评估、临床基准和通用推理研究,用于评估大语言模型对临床案例回答中的表达性推理,旨在使评估决策透明化。

AI 中文摘要

评分量规支持对语言模型进行结构化评估。我们提出了一种用于评估模型回答中表达性临床推理的评分量规,该量规借鉴了三方面的工作:医学教育评估框架(ART、SCT、关键特征问题和OSCE);临床大语言模型基准(MedR-Bench、HealthBench、TIMER-Bench、this http URL、PrIME-LLM和PatientSafeBench);以及通用大语言模型推理评估研究,包括事实性-有效性-连贯性-实用性分类法、FaithCoT-Bench和C2-Faith。我们将接地性(groundedness)作为该分类法事实性类别的一种临床导向改编。该评分量规将这些概念整合到一个多维框架中,用于对黄金标准临床案例的开放式回答进行评分。它包含暂定的行为锚点、适用规则以及针对特定案例的安全关键性错误的单独标记。通用领域框架为其设计提供了参考,但并未被视为经过验证的临床工具。该评分量规不能替代特定案例的参考标准或现有基准的任务特定指标。它尚未经过评分者间信度、结构效度或临床实用性的测试。其直接目的是在实证测试之前,使评估决策明确化并接受审查。

英文摘要

We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.

Comments20 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑