arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MyoCardBench:用于评估临床真实心血管护理场景中大型语言模型的真实世界数据基准

CardioBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

Xiao Li, Mouxiao Bian, Zhaodi Wu, Sijie Ren, Juechen Chen, Lu Lu, Jingru Ding, Yun Zhong, Jie Xu, Yixiu Liang, Junbo Ge

arXiv 2607.25186首次发表:更新:

发表机构

Shanghai Institute of Cardiovascular Diseases; Shanghai Artificial Intelligence Laboratory; Minhang Hospital, Fudan University(上海市心血管病研究所; 上海人工智能实验室; 复旦大学附属闵行医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究开发MyoCardBench真实世界基准评估大语言模型在心血管护理场景的性能,涵盖多任务数据集,经专家注释审核。7个模型在零样本设置下输出,GPT-5.4表现最佳。此基准为评估模型提供框架,助于发现优势与不足。

AI 中文摘要

背景:大多数医学大语言模型基准专注于检查知识或孤立任务,可能无法反映心血管护理的纵向、多模态和安全关键工作流程。目的:开发MyoCardBench,一个跨越心血管护理连续过程的真实世界基准,并评估大语言模型在临床维度和专业任务上的性能。方法:MyoCardBench包含来自13个特定任务数据集的2263个项目,这些数据集来自去识别化的心血管记录和检查数据。16位心脏病专家进行注释和参考构建,随后由两位高级心脏病专家进行交叉审核。7个大语言模型在标准化零样本设置下生成了15841个输出。开放式任务使用关键点覆盖率和整体临床质量进行评估,而CardioEthics通过准确率进行评分。结果:GPT-5.4获得最高的宏观平均(62.55)和项目加权均值(62.19),其次是Gemini 3.1 Pro(59.95)和Qwen 3.6 27B(59.72)。GPT-5.4在所有三个维度上均排名第一。CardioAuxReport表现最佳(86.38),而CardioECGRead(17.25)和CardioEthics(17.34)最低。整体临床质量和关键点覆盖率之间最大的差距出现在CardioComm(52.71)、CardioEmergRescue(52.05)和CardioTreatPlan(48.80)中。结论:据我们所知,MyoCardBench是用于评估心血管护理连续过程中大型语言模型的最大的真实世界、多任务基准,并且提供了迄今为止报道的最广泛的临床真实心脏病学场景覆盖范围。它为识别模型优势、临床上重要的遗漏以及未来发展的优先事项提供了一个严格的框架。

英文摘要

Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop CardioBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: CardioBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-weighted mean (62.19), followed by Gemini 3.1 Pro (59.95) and Qwen 3.6 27B (59.72). GPT-5.4 ranked first in all three dimensions. CardioAuxReport performed best (86.38), whereas CardioECGRead (17.25) and CardioEthics (17.34) were lowest. The largest gaps between holistic clinical quality and key-point coverage occurred in CardioComm (52.71), CardioEmergRescue (52.05), and CardioTreatPlan (48.80). Conclusions: To our knowledge, CardioBench is the largest real-world, multi-task benchmark for LLM evaluation across the cardiovascular care continuum and offers the broadest coverage of clinically authentic cardiology scenarios reported to date. It provides a rigorous framework for identifying model strengths, clinically important omissions, and priorities for future development.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑