arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CDGC-Net:基于协作双尺度自注意力与分组通道建模的3D医学图像分割网络

CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling

Zheyang Jing, Qin Lu, Jianwang Li, Yujie Yang, Chen Yi, Shaofeng Jiang

arXiv 2608.08575首次发表:更新:

发表机构

Nanchang Hangkong University(南昌航空大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出CDGC-Net,通过协作双尺度自注意力与分组通道建模优化3D医学图像分割,在多数据集上精度优于现有方法,且参数量与计算量较UNETR++显著降低,实现了精度与复杂度的良好平衡。

AI 中文摘要

准确的3D医学图像分割需要将长程解剖上下文与精细边界细节相结合。现有方法常将全局和局部特征在独立模块或特征层级中建模,并单独进行通道重校准,这可能导致全局上下文与局部边界间的语义不匹配、通道关系建模不足、空间-通道交互薄弱以及表示冗余。我们提出CDGC-Net,一种结合协作双尺度空间注意力与分组分层通道建模的3D医学图像分割网络。在每个CDGC模块内,协作双尺度自注意力(CDSA)将注意力头分配到并行的局部窗口分支和全局稀疏分支,两个分支在同一特征层级同时捕获精细空间细节与长程解剖上下文,其输出被拼接为N×C的空间表示并直接传递给分组分层通道注意力(GHCA);GHCA将通道组织为r组,同时建模组内和组间依赖关系;CDSA与GHCA复用共享的键投影以维持一致的特征参考,随后通过残差特征对齐将优化后的特征与原始表示融合。在Synapse、ACDC、BraTS和LA数据集上,CDGC-Net分别取得86.96%、92.91%、82.56%和93.52%的平均DSC值,较次高报告值分别超出0.39、0.47、0.17和0.32个百分点;对于64×128×128的输入尺寸,CDGC-Net包含25.83M参数和28.62G FLOPs,较UNETR++分别减少39.87%和40.30%,这些结果表明其在分割精度与计算复杂度间实现了良好的权衡。

英文摘要

Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often model global and local features in separate modules or feature levels and perform channel recalibration independently. This may cause semantic mismatch between global context and local boundaries, insufficient channel relationship modeling, weak spatial-channel interaction, and redundant representations. We propose CDGC-Net, a 3D medical image segmentation network that combines cooperative dual-scale spatial attention with grouped hierarchical channel modeling. With-in each CDGC block, Cooperative Dual-Scale Self-Attention (CDSA) assigns attention heads to parallel local-window and global-sparse branches. The two branches capture fine spatial details and long-range anatomical context at the same feature level. Their outputs are concatenated into an $N\times C$ spatial representation and directly passed to Grouped Hierarchical Channel Attention (GHCA). GHCA organizes the channels into $r$ groups and models both within-group and cross-group dependencies. CDSA and GHCA reuse a shared key projection to maintain a consistent feature reference. Residual feature alignment subsequently integrates the refined features with the original representation. On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values of 86.96\%, 92.91\%, 82.56\%, and 93.52\%, respectively, exceeding the next-highest reported values by 0.39, 0.47, 0.17, and 0.32 percentage points. CDGC-Net contains 25.83M parameters and 28.62G FLOPs for an input size of $64\times128\times128$, reducing these quantities by 39.87\% and 40.30\%, respectively, relative to UNETR++. These results indicate a favorable trade-off between segmentation accuracy and computational complexity.

Journal ref14th International Conference on Image and Graphics (ICIG 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑