TomoTransformer:面向CT重建的基础模型
TomoTransformer: Towards a Foundation Model for CT Reconstruction
- Swiss Data Science Center (SDSC) in Paul Scherrer Institute (PSI)(瑞士数据科学中心(SDSC),保罗·谢勒研究所(PSI))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出TomoTransformer,一种基于Transformer的CT重建基础模型,在反投影空间中对局部投影进行自注意力处理,实现任意数量输入输出投影的无需重训练泛化,并在基准和真实同步辐射数据上表现优异。
AI中文摘要:
监督式深度学习推动了稀疏视角断层重建的发展。然而,传统模型通常将滤波反投影(FBP)图像或正弦图映射到干净的重建结果,这类模型在分布偏移下表现脆弱。由于每当投影数量和角度、探测器分辨率或数据分布发生变化时都需要重新训练,它们在实际应用中的部署仍然受限。为解决这一问题,我们提出了TomoTransformer,一种基于Transformer的架构,它将每个局部滤波投影视为独立的令牌,并通过自注意力机制预测缺失视角。关键在于,TomoTransformer在反投影空间中运行,该空间将投影按空间位置分离,使得视角插值在几何上良定且对探测器尺寸不变。这一设计产生了一个单一的基础模型,能够处理任意数量的输入投影,在任意角度位置和探测器维度下工作,并无需重新训练即可查询任意数量的目标角度。在一个涵盖多种医学CT解剖结构和自然图像的大规模数据集上训练后,TomoTransformer在解剖结构、材料和分辨率上均能有效泛化。在多个基准稀疏视角数据集上的广泛评估表明,TomoTransformer显著优于如ViewTrans等同期多用途模型,并达到或超过强协议特定基线,同时完全不受输入和目标投影数量的限制。此外,该模型在从X射线同步辐射源收集的真实实验纳米级脑数据上展示了鲁棒的零样本泛化能力,彰显了其在实际应用中的实用价值。
英文摘要:
Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.