arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04514cs.CL

RESPClinBench:呼吸专科医疗中的多模态临床决策与纵向疾病管理基准测试

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Mouxiao Bian, Zhi Chen, Ruiyao Chen, Lu Lu, Hengrui Liang, Chaoyi Huang, Yiluo Lin, Jingru Ding, Yun Zhong, Yueming Su, Jie Xu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究开发了呼吸专科多模态临床决策与纵向疾病管理基准RESPClinBench,在两个数据集上评估7种大模型,识别模型在肺结节评估和COPD管理的局限,为模型选择提供临床依据。

中文摘要 AI 辅助

背景:呼吸专科医疗需要多模态解读、纵向风险评估、符合指南的干预措施及全程管理,而当前以检查为导向的医疗基准未能充分体现这些需求。目的:开发基于真实场景的呼吸临床决策基准RESPClinBench,并在AECOPD-PIM和PNBIM两个数据集上评估7种现有大型语言模型。方法:RESPClinBench的案例改编自去标识化的呼吸临床数据,3名主治级呼吸科医生对案例、参考回答及原子临床操作要点进行修订,1名资深呼吸科专家完成交叉审核与最终裁定。其中AECOPD-PIM包含427例开放性慢性阻塞性肺疾病(COPD)案例,PNBIM包含196例结合胸部CT与结构化临床信息的多模态肺结节案例;7种模型通过标准化API推理生成4361条响应,推理时温度设置为0,最大输出长度为8192个token,采用自动化框架计算最终得分,得分由原子操作召回率与基于评分规则的大模型作为评判者(LLM-as-a-Judge)评估结果的算术平均值得出。结果:在全部623例案例中,平均最终得分为68.58;Qwen3.6-27B以71.22的总分位列第一,Qwen3.5-397B-A17B在PNBIM数据集上以72.48的得分领先,Qwen3.6-27B在AECOPD-PIM数据集上以71.11的得分领先;PNBIM响应中影像幻觉和严重医疗风险的占比分别为31.85%和8.16%,AECOPD-PIM响应中用药安全风险和严重医疗风险的占比分别为26.93%和1.44%。结论:RESPClinBench可识别多模态肺结节评估及纵向COPD管理中的任务特定局限性,结合明确的临床操作覆盖、整体评估及独立安全标记,为模型选择与前瞻性验证提供了基于临床实际的依据。

英文摘要

Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course management, which are poorly represented by examination-oriented medical benchmarks. Objective: To develop RESPClinBench, a real-world scenario-based benchmark for respiratory clinical decision-making, and evaluate seven contemporary large language models across AECOPD-PIM and PNBIM. Methods: RESPClinBench cases were adapted from de-identified respiratory clinical data. Three attending-level respiratory physicians revised cases, reference answers, and atomic clinical-action points, while one senior respiratory specialist performed cross-review and final adjudication. AECOPD-PIM comprised 427 open-ended COPD cases, and PNBIM comprised 196 multimodal pulmonary nodule cases combining chest CT with structured clinical information. Seven models generated 4,361 responses through standardized API inference with temperature 0 and a maximum output length of 8192 tokens. An automated framework calculated the final score as the arithmetic mean of atomic-action recall and rubric-based LLM-as-a-Judge assessment. Results: Across 623 cases, the mean final score was 68.58. Qwen3.6-27B ranked first overall at 71.22, Qwen3.5-397B-A17B led PNBIM at 72.48, and Qwen3.6-27B led AECOPD-PIM at 71.11. Imaging hallucination and serious medical risk occurred in 31.85% and 8.16% of PNBIM responses; medication-safety risk and serious medical risk occurred in 26.93% and 1.44% of AECOPD-PIM responses. Conclusions: RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management. Combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.

发表机构

  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Macau University of Science and Technology(澳门科技大学)
  • First Affiliated Hospital of Guangzhou Medical University(广州医科大学附属第一医院)
  • Guangzhou Institute of Respiratory Health(广州呼吸健康研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑