发表机构
Khalifa University(哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MedCORE提出一种基于临床标准分解、多尺度证据编码和图注意力依赖建模的视觉-语言诊断框架,在三个医学数据集上超越多种强基线,提升可解释性与诊断性能。
AI 中文摘要
临床诊断本质上是一个结构化的推理过程,然而现有的深度学习模型往往绕过这一结构,直接将图像特征映射到疾病标签,而未能明确检验临床医生系统评估的形态学和纹理标准。这限制了诊断的透明度,并可能危及安全临床部署。我们提出MedCORE(医学标准导向推理与证据),一种结构化诊断框架,在视觉-语言架构中实现临床推理的操作化。对于每张输入图像,MedCORE将诊断过程分解为临床定义的标准,将每个标准空间定位到诊断相关的图像区域,通过捕捉宏观结构和微观纹理病理特征的多尺度表示编码证据,并使用图注意力网络细化标准表示,该网络显式建模标准间依赖关系。标准表示进一步与临床文本描述符对齐,通过类别级视觉原型强化,并使用不确定性校准的加权进行聚合,该加权按比例降低低置信度诊断证据的权重。MedCORE在三种临床异质性成像模态上得到验证,包括ISIC 2018上的皮肤镜病变分类、BUSI上的乳腺超声病变表征以及IDRiD上的糖尿病视网膜病变分级。定量上,MedCORE在ISIC 2018上达到89.2%的准确率、85.7%的宏F1和96.4%的AUC;在BUSI上达到96.1%的准确率、95.2%的宏F1和98.4%的AUC;在IDRiD上达到84.3%的准确率、80.2%的宏F1和92.8%的AUC。这些结果表明,相对于强CNN、Transformer、生物医学视觉-语言、基于概念和基于原型的基线,MedCORE均取得了一致的改进。
英文摘要
Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image features directly to disease labels without explicitly interrogating the morphological and textural criteria that clinicians systematically evaluate. This limits diagnostic transparency and may compromise safe clinical deployment. We present MedCORE (Medical Criteria-Oriented Reasoning and Evidence), a structured diagnostic framework that operationalizes clinical reasoning within a vision-language architecture. For each input image, MedCORE decomposes the diagnostic process into clinically defined criteria, spatially localizes each criterion to diagnostically relevant image regions, encodes evidence through multi-scale representations that capture macro-structural and micro-textural pathological characteristics, and refines criterion representations using a Graph Attention Network that explicitly models inter-criteria dependencies. Criterion representations are further aligned with clinical text descriptors, reinforced through class-wise visual prototypes, and aggregated using uncertainty-calibrated weighting that proportionally discounts low-confidence diagnostic evidence. MedCORE is validated across three clinically heterogeneous imaging modalities, including dermoscopic lesion classification on ISIC 2018, breast ultrasound lesion characterization on BUSI, and diabetic retinopathy grading on IDRiD. Quantitatively, MedCORE achieves 89.2% accuracy, 85.7% macro-F1, and 96.4% AUC on ISIC 2018; 96.1% accuracy, 95.2% macro-F1, and 98.4% AUC on BUSI; and 84.3% accuracy, 80.2% macro-F1, and 92.8% AUC on IDRiD. These results demonstrate consistent improvements over strong CNN, transformer, biomedical vision-language, concept-based, and prototype-based baselines.
Comments16 pages, 4 figures, conference