Teeth2Point:一种两阶段牙科CBCT ROI到点的分割框架
Teeth2Point: A Two-Stage Dental CBCT ROI-to-Point Segmentation Framework
- ETH Zurich(苏黎世联邦理工学院)
- Align Technology(Align Technology公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对牙科CBCT中牙齿缺失或错位的标注难题,提出Teeth2Point框架,通过卷积定位ROI、自适应采样转点令牌,结合SSL预训练与监督微调,在四个数据集上较基线显著提升异常病例分割性能。
AI中文摘要:
现代深度学习架构在牙科CBCT分割中已展现出强大性能。一个仍存在的关键挑战是在牙齿缺失或错位的病例中实现准确的牙齿标注,这与牙科实践高度相关。基于Transformer的架构理论上应能利用全局解剖上下文解决此类歧义。然而,由于CBCT体积的高分辨率以及牙齿在体积内的广泛空间分布,基于密集块的体积处理面临固有的权衡:计算成本限制了自注意力中可使用的块数量,因此要么增加自注意力中捕获的上下文范围,要么通过使用小块捕获细粒度结构细节,但无法同时实现两者。本研究提出Teeth2Point,一种用于牙科CBCT语义分割的高效基于点的Transformer框架,可避免该权衡。Teeth2Point首先使用卷积模型定位牙齿周围的体积感兴趣区域(ROI),随后通过自适应采样将ROI转换为点令牌。Transformer模型利用这些点令牌预测准确的分割结果,既能捕获全局上下文又能保持高分辨率。该Transformer首先以自监督学习(SSL)方式进行预训练,采用DINO的风格但使用特定领域的增强策略,随后进行监督微调。包含随机令牌掩码的SSL预训练可提升对复杂解剖变异的鲁棒性。与最强的两阶段基线相比,Teeth2Point在四个数据集上的异常病例性能平均提升1.44 DSC点;与第一阶段nnU-Net相比,提升了1.9个点。
英文摘要:
Modern deep learning architectures have demonstrated strong performance in dental CBCT segmentation. One remaining crucial challenge is accurate tooth labeling in cases with missing or malpositioned teeth, which are highly relevant for dental practice. Transformer-based architectures should in theory be able to resolve such ambiguities using global anatomical context. However, due to the high resolution of CBCT volumes and the wide spatial distribution of teeth within volumes, dense patch-based volumetric processing faces an inherent trade-off. Computational costs limit the number of patches that can be used in self-attention and thus, one can either increase the extent of the context captured in self-attention or capture fine-grained structural details by using small patches, but not both. In this work, we present Teeth2Point, an efficient point-based transformer framework for dental CBCT semantic segmentation that can avoid this trade-off. Teeth2Point first localizes volumetric regions of interest (ROIs) surrounding teeth using a convolutional model, then converts ROIs into point tokens using adaptive sampling. A transformer model predicts accurate segmentations using the point tokens, which allow capturing global context while retaining high resolution. The transformer is first pretrained using self-supervised learning (SSL), in the style of DINO but using domain-specific augmentation strategies, followed by supervised finetuning. The SSL pretraining, which includes random token masking, provides robustness to complex anatomical variations. Compared with the strongest two-stage baseline, Teeth2Point improves abnormal-case performance by 1.44 DSC points on average across four datasets; relative to the first-stage nnU-Net, the gain is 1.9 points.