arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15888cs.CVcs.AI

解剖学锚定与泄漏感知的多模态对比学习用于基于结构MRI的阿尔茨海默病分类

Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI

Paul-Gabriel Nicolae, Irina Georgiana Mocanu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对AD分类中模型关注无关区域及标签泄漏问题,提出解剖学锚定与泄漏感知的对比学习框架,并通过内侧颞叶裁剪将图像准确率提升至65.1%。

中文摘要 AI 辅助

用于阿尔茨海默病(AD)分期诊断的结构MRI深度网络,在达到合理准确率的同时,常常关注与解剖学无关的区域;而加入临床表格的多模态模型,则经常依赖那些最初用于分配诊断标签的变量。我们使用一个刻意轻量化的基于切片的编码器(ResNet18配合一层Transformer处理切片),在ADNI-1的1,075例基线T1加权扫描上研究这两个问题。首先,我们使用FastSurfer分割结果作为解剖学参考:基于分割标签训练的YOLOv8模型定位AD相关结构的mAP_50超过0.96,而Grad-CAM对比显示,仅图像分类器经常关注颅骨、眼眶和背景区域。其次,我们调整了CLIP风格的图像-表格对比学习框架,并将ADNIMERGE变量按照标签泄漏谱进行组织。与认知评分融合后,三分类准确率达到87.3%,我们将其视为由泄漏驱动的上限,而非成像结果;与区域体积融合后,准确率为73.0%。我们观察到,对比目标的选择会改变图像编码器学习的内容:在MCI与CN的分类中,当编码器与认知评分对齐时,仅图像头的准确率为52.4%;当与体积对齐时,准确率为73.8%,尽管推理时未使用任何表格输入。第三,将输入限制为每个受试者内侧颞叶的裁剪区域,仅图像的三分类准确率从58.7%提升至65.1%。所有结果均来自小型平衡测试集上的单次运行,我们报告了置信区间以及阻止与已发表数字直接比较的协议差异。

英文摘要

Deep networks trained on structural MRI for Alzheimer's disease (AD) staging often reach reasonable accuracy while attending to anatomically irrelevant regions, and multimodal models that add clinical tables frequently rely on variables that were used to assign the diagnostic label in the first place. We study both issues with a deliberately lightweight slice-based encoder (ResNet18 with a one-layer Transformer over slices) on 1,075 baseline T1-weighted scans from ADNI-1. First, we use FastSurfer segmentations as an anatomical reference: YOLOv8 models trained on segmentation-derived labels localize Alzheimer-relevant structures with mAP_50 above 0.96, and a Grad-CAM comparison shows that the image-only classifier frequently attends to the skull, orbits and background. Second, we adapt a CLIP-style image - tabular contrastive framework and organize ADNIMERGE variables along a label-leakage spectrum. Fusion with cognitive scores yields 87.3% three-way accuracy, which we treat as a leakage-driven upper bound rather than an imaging result; fusion with regional volumes yields 73.0%. We observe that the choice of contrastive target changes what the image encoder learns: on MCI vs. CN, the image-only head reaches 52.4% when the encoder is aligned to cognitive scores and 73.8\% when aligned to volumes, although no tabular input is used at inference. Third, restricting the input to a per-subject crop of the medial temporal lobe raises image-only three-way accuracy from 58.7% to 65.1%. All results come from single runs on a small balanced test set, and we report confidence intervals and the protocol differences that prevent direct comparison with published numbers.

发表机构

  • National University of Science and Technology POLITEHNICA Bucharest(布加勒斯特国立科技理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑