在U-Net变体中利用几何先验进行多聚体分割
Learning with Geometric Priors in U-Net Variants for Polyp Segmentation
- The University of Texas Rio Grande Valley(德克萨斯理工大学里奥格兰德分校)
- Sewickley Academy(塞维克利学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Geometric Prior-guided Module,通过引入几何先验提升U-Net变体在多聚体分割中的性能。
AI中文摘要:
准确且稳健的多聚体分割对于早期结直肠癌检测和计算机辅助诊断至关重要。尽管基于卷积神经网络、Transformer和Mamba的U-Net变体已取得了强劲性能,但它们仍然难以捕捉几何和结构线索,尤其是在低对比或杂乱的内窥镜场景中。为了解决这一挑战,我们提出了一种新颖的几何先验引导模块(GPM),该模块将显式的几何先验注入到基于U-Net的架构中,用于多聚体分割。具体而言,我们对视觉几何基础Transformer(VGGT)在模拟的ColonDepth数据集上进行微调,以估计适合内窥镜领域的多聚体图像的深度图。这些深度图随后由GPM处理,将几何先验编码到编码器的特征图中,其中它们进一步通过空间和通道注意力机制进行细化,强调局部空间和全局通道信息。GPM是即插即用的,可以无缝集成到多种U-Net变体中。在五个公开的多聚体分割数据集上进行的广泛实验显示,与三个强大的基线相比,取得了持续的改进。代码和生成的深度图可在:https://github.com/fvazqu/GPM-PolypSeg上获得。
英文摘要:
Accurate and robust polyp segmentation is essential for early colorectal cancer detection and for computer-aided diagnosis. While convolutional neural network-, Transformer-, and Mamba-based U-Net variants have achieved strong performance, they still struggle to capture geometric and structural cues, especially in low-contrast or cluttered colonoscopy scenes. To address this challenge, we propose a novel Geometric Prior-guided Module (GPM) that injects explicit geometric priors into U-Net-based architectures for polyp segmentation. Specifically, we fine-tune the Visual Geometry Grounded Transformer (VGGT) on a simulated ColonDepth dataset to estimate depth maps of polyp images tailored to the endoscopic domain. These depth maps are then processed by GPM to encode geometric priors into the encoder's feature maps, where they are further refined using spatial and channel attention mechanisms that emphasize both local spatial and global channel information. GPM is plug-and-play and can be seamlessly integrated into diverse U-Net variants. Extensive experiments on five public polyp segmentation datasets demonstrate consistent gains over three strong baselines. Code and the generated depth maps are available at: https://github.com/fvazqu/GPM-PolypSeg