AI 中文总结
ProgFormer是一种分层体素扩散Transformer,通过粗-细通路结合条件流匹配,在ADNI等三个基准上实现了更优的纵向脑部MRI预测性能。
AI 中文摘要
预测脑部未来结构MRI具有挑战性,因为纵向变化通常较为细微且局限于特定解剖区域,而大多数特定受试者的脑部结构随时间保持稳定。因此,有效模型应在保留全局脑部结构一致性的同时,对细粒度疾病进展保持敏感。现有基于隐空间的方法提高了计算效率,但在压缩-重建过程中存在信息损失;相比之下,直接体素空间方法虽避免了隐式重建,但通常使用统一预测通路来建模脑部结构和进展相关变化,细微局部变化可能被占主导的稳定脑部结构所掩盖。为应对这些挑战,我们提出ProgFormer,一种用于纵向脑部MRI预测的分层体素空间扩散Transformer。ProgFormer使用粗通路从3D块标记执行初级体积预测,该通路建模整体脑部结构和纵向上下文;细通路则以粗表示作为时空基础,在单个块内进行体素级细化。两条通路通过条件流匹配直接在体素空间中联合估计速度场,实现端到端预测,无需单独学习的图像自动编码器。随后通过欧拉步序列积分估计的速度场,从高斯噪声生成预测的未来扫描。在ADNI、AIBL和OASIS三个广泛使用的基准上,分别在成对和轨迹设置下进行的大量实验结果表明,与几种最先进方法相比,ProgFormer表现出良好性能。
英文摘要
Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, while most subject-specific brain structure remains stable over time. An effective model should therefore preserve global brain structural consistency while remaining sensitive to fine-grained disease progression. Existing latent-space-based methods improve computational efficiency, but suffer from information loss during their compression-reconstruction procedure. In contrast, direct voxel-space methods avoid latent reconstruction but commonly use a unified prediction pathway to model brain structure and progression-related changes. Subtle local changes may therefore be overshadowed by the dominant stable brain structure. To address these challenges, we propose ProgFormer, a hierarchical voxel-space Diffusion Transformer for longitudinal brain MRI prediction. ProgFormer uses a coarse pathway to perform the primary volumetric prediction from 3D patch tokens. This pathway models overall brain structure and longitudinal context. The fine pathway then uses the coarse representations as spatio-temporal grounding for voxel-level refinement within individual patches. The two pathways jointly estimate a velocity field directly in voxel space through conditional flow matching, enabling end-to-end prediction without a separately learned image autoencoder. The predicted future scan is then generated from Gaussian noise by integrating the estimated velocity field over a sequence of Euler steps. Extensive experimental results on three widely used benchmarks, ADNI, AIBL, and OASIS, under both pairwise and trajectory settings demonstrate favourable performance compared against several state-of-the-art methods.