发表机构
Osnabrück University; Rhodes University; National Institute for Theoretical and Computational Sciences (NITheCS); Madanapalle Institute of Technology and Science; Université Clermont Auvergne(奥斯纳布吕克大学; 罗德斯大学; 国家理论与计算科学研究所; 马达纳帕莱理工学院; 克莱蒙奥弗涅大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出可解释多模态深度学习框架,结合3D MRI与临床数据诊断阿尔茨海默病,发现性能与解释依赖任务、模态、融合及队列,跨队列需评估可解释性。
AI 中文摘要
痴呆症是一个重大且日益增长的全球健康负担,其中阿尔茨海默病(AD)占大多数病例。及时准确的诊断对于管理这一负担至关重要,并且日益依赖于整合互补的临床和影像信息。多模态深度学习可以结合这些模态进行AD诊断,但其解释在不同模态、融合策略和队列中的表现仍不清楚。我们开发了一个可解释的多模态框架,将用于T1加权MRI的3D CNN编码器与用于协调临床和人口统计学数据的前馈网络配对,在ADNI的6,479条内部记录和OASIS-3的1,703条独立记录上,对三分类和两两分类诊断任务比较了多种模型设置。在ADNI上,仅表格模型实现了最高的三分类AUC-ROC为0.879,并在区分认知正常(CN)与轻度认知障碍(MCI)方面表现最佳(0.903),而交叉注意力在MCI与AD区分上表现最好(0.861);CN与AD整体上具有高度判别性。在OASIS-3上,仅视觉模型表现最佳(三分类AUC-ROC为0.910);CN与MCI的区分仍然困难,且没有任何融合策略在任务和队列中始终优于单一模态。SHAP和Integrated Gradients识别出MMSE在两个队列中都是主导的表格特征,全局特征排名在ADNI(ρ=0.94)和OASIS-3(ρ=0.96)中高度一致;然而,基于CAM的解释随模型配置和队列而变化。这些发现表明,多模态性能和解释依赖于任务、模态、融合和队列:一个主导的认知信号在队列间持续存在,但特征贡献和CAM解释并非如此,这强调了在队列偏移下评估可解释性的必要性,而非将其视为稳定的内在属性。
英文摘要
Dementia is a major and growing global health burden, with Alzheimer's disease (AD) accounting for most cases. Timely and accurate diagnosis is central to managing this burden and increasingly depends on integrating complementary clinical and imaging information. Multimodal deep learning can combine these modalities for AD diagnosis, but how its explanations behave across modalities, fusion strategies, and cohorts remains unclear. We developed an explainable multimodal framework pairing a 3D CNN encoder for T1-weighted MRI with a feedforward network for harmonized clinical and demographic data, comparing varied model setups on three-way and pairwise diagnostic tasks using 6,479 internal records from the ADNI and 1,703 independent records from the OASIS-3. On ADNI, the tabular-only model achieved the highest three-class AUC-ROC of 0.879 and best discriminated cognitively normal (CN) versus mild cognitive impairment (MCI; 0.903), while cross-attention performed best for MCI versus AD (0.861); CN versus AD was highly discriminative overall. On OASIS-3, the vision-only model performed best (three-class AUC-ROC 0.910); CN versus MCI remained difficult, and no fusion strategy consistently outperformed single modalities across tasks and cohorts. SHAP and Integrated Gradients identified the MMSE as the dominant tabular feature in both cohorts, with global feature rankings agreeing strongly in ADNI ($ρ=0.94$) and OASIS-3 ($ρ=0.96$); CAM-based explanations, however, changed with model configuration and cohort. These findings show that multimodal performance and explanations are task, modality, fusion, and cohort-dependent: a dominant cognitive signal persisted across cohorts, but feature contributions and CAM explanations did not, underscoring the need to evaluate explainability under cohort shift rather than as a stable, intrinsic property.