M$^2$PFN:用于阿尔茨海默病泛化多模态上下文学习的端到端解耦对齐
M$^2$PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease
浏览论文内容
中文总结 AI 辅助
M$^2$PFN通过端到端解耦对齐将表格基础模型TabPFN扩展为多模态AD预测器,在ADNI及外部队列上取得最优性能。
中文摘要 AI 辅助
尽管已有多种结合影像和表格数据的多模态方法用于阿尔茨海默病(AD)诊断,但这些方法在跨队列泛化方面往往受限。上下文学习(ICL)在诸如TabPFN等基础表格模型中展现了优异的泛化性能和高度灵活性。为了将TabPFN的ICL扩展到多模态AD分析,主要障碍在于TabPFN是在合成表格先验上进行元训练的,这些先验与图像衍生特征的统计结构并不自然匹配。我们提出M$^2$PFN,一个端到端框架,将该表格基础模型转变为多模态AD预测器。M$^2$PFN(i)通过TabPFN的Transformer进行可微推理,将任务梯度反向传播到3D-MRI和表格编码器;(ii)通过解耦和对比目标,将两种模态对齐到与ICL引擎先验匹配的共享子空间;(iii)通过可学习的门控捷径整合冻结的纯表格预测。由于ICL引擎保持冻结,其上下文机制得以保留以进行测试时泛化,而端到端训练则将编码器塑造成其可利用的特征。在ADNI($n=2240$,三分类CN/MCI/AD)上,M$^2$PFN达到$65.55\%$的宏F1和$82.21\%$的宏AUC,超越了全面的单模态和多模态基线。仅将头部替换为TabPFN回归器,同一架构在$1250$名受试者的子队列上回归基线MMSE,测试MAE为$1.743$,优于所有多模态基线。在两个外部队列(OASIS-3和SCAN)上,无需重新训练,M$^2$PFN在所有基线中取得了最佳AUC和最低MMSE MAE,并且即使在认知工具改变时也能迁移。
英文摘要
While various multimodal methods combining imaging and tabular data for Alzheimer's disease (AD) diagnosis were proposed, they are often limited in generalization across cohorts. In-context learning (ICL) has demonstrated excellent generalization performances and high flexibility in foundational tabular models such as TabPFN. To extend TabPFN's ICL to multimodal AD analysis, the main obstacle is that TabPFN is meta-trained on synthetic tabular priors that do not naturally match the statistical structure of image-derived features. We propose M$^2$PFN, an end-to-end framework that turns this tabular foundation model into a multimodal AD predictor. M$^2$PFN (i) performs differentiable inference through TabPFN's transformer, back-propagating task gradients into 3D-MRI and tabular encoders; (ii) aligns the two modalities into a shared subspace, via disentanglement and a contrastive objective, matched to the ICL engine's prior; and (iii) folds in a frozen tabular-only prediction through a learnable gated shortcut. Because the ICL engine stays frozen, its in-context mechanism is preserved for test-time generalization, while end-to-end training shapes the encoders into features it can exploit. On ADNI ($n=2240$, three-class CN/MCI/AD), M$^2$PFN attains $65.55\%$ macro-F1 and $82.21\%$ macro-AUC, surpassing a comprehensive set of unimodal and multimodal baselines. By swapping only the head for a TabPFN regressor, the same architecture regresses baseline MMSE on a $1250$-subject sub-cohort to test MAE $1.743$, outperforming every multimodal baseline. On two external cohorts (OASIS-3 and SCAN) with no retraining, M$^2$PFN achieves the best AUC and the lowest MMSE MAE across all baselines, and transfers even when the cognitive instrument changes.
发表机构
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。