RESPClinBench:呼吸专科医疗中的多模态临床决策与纵向疾病管理基准测试
RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care
浏览论文内容
中文总结 AI 辅助
该研究开发了呼吸专科多模态临床决策与纵向疾病管理基准RESPClinBench,在两个数据集上评估7种大模型,识别模型在肺结节评估和COPD管理的局限,为模型选择提供临床依据。
中文摘要 AI 辅助
背景:呼吸专科医疗需要多模态解读、纵向风险评估、符合指南的干预措施及全程管理,而当前以检查为导向的医疗基准未能充分体现这些需求。目的:开发基于真实场景的呼吸临床决策基准RESPClinBench,并在AECOPD-PIM和PNBIM两个数据集上评估7种现有大型语言模型。方法:RESPClinBench的案例改编自去标识化的呼吸临床数据,3名主治级呼吸科医生对案例、参考回答及原子临床操作要点进行修订,1名资深呼吸科专家完成交叉审核与最终裁定。其中AECOPD-PIM包含427例开放性慢性阻塞性肺疾病(COPD)案例,PNBIM包含196例结合胸部CT与结构化临床信息的多模态肺结节案例;7种模型通过标准化API推理生成4361条响应,推理时温度设置为0,最大输出长度为8192个token,采用自动化框架计算最终得分,得分由原子操作召回率与基于评分规则的大模型作为评判者(LLM-as-a-Judge)评估结果的算术平均值得出。结果:在全部623例案例中,平均最终得分为68.58;Qwen3.6-27B以71.22的总分位列第一,Qwen3.5-397B-A17B在PNBIM数据集上以72.48的得分领先,Qwen3.6-27B在AECOPD-PIM数据集上以71.11的得分领先;PNBIM响应中影像幻觉和严重医疗风险的占比分别为31.85%和8.16%,AECOPD-PIM响应中用药安全风险和严重医疗风险的占比分别为26.93%和1.44%。结论:RESPClinBench可识别多模态肺结节评估及纵向COPD管理中的任务特定局限性,结合明确的临床操作覆盖、整体评估及独立安全标记,为模型选择与前瞻性验证提供了基于临床实际的依据。
英文摘要
Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course management, which are poorly represented by examination-oriented medical benchmarks. Objective: To develop RESPClinBench, a real-world scenario-based benchmark for respiratory clinical decision-making, and evaluate seven contemporary large language models across AECOPD-PIM and PNBIM. Methods: RESPClinBench cases were adapted from de-identified respiratory clinical data. Three attending-level respiratory physicians revised cases, reference answers, and atomic clinical-action points, while one senior respiratory specialist performed cross-review and final adjudication. AECOPD-PIM comprised 427 open-ended COPD cases, and PNBIM comprised 196 multimodal pulmonary nodule cases combining chest CT with structured clinical information. Seven models generated 4,361 responses through standardized API inference with temperature 0 and a maximum output length of 8192 tokens. An automated framework calculated the final score as the arithmetic mean of atomic-action recall and rubric-based LLM-as-a-Judge assessment. Results: Across 623 cases, the mean final score was 68.58. Qwen3.6-27B ranked first overall at 71.22, Qwen3.5-397B-A17B led PNBIM at 72.48, and Qwen3.6-27B led AECOPD-PIM at 71.11. Imaging hallucination and serious medical risk occurred in 31.85% and 8.16% of PNBIM responses; medication-safety risk and serious medical risk occurred in 26.93% and 1.44% of AECOPD-PIM responses. Conclusions: RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management. Combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.
发表机构
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Macau University of Science and Technology(澳门科技大学)
- First Affiliated Hospital of Guangzhou Medical University(广州医科大学附属第一医院)
- Guangzhou Institute of Respiratory Health(广州呼吸健康研究院)
机构由 AI 辅助整理,请以论文原文为准。