发表机构
Tamkang University(淡江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MSCA-UNet,在U-Net瓶颈处结合多尺度空洞卷积上下文聚合与解码器CBAM注意力细化,将mIoU从基线96.9%提升至99.1%,验证了两者的互补增益。
AI 中文摘要
U-Net因其简单的编码器-解码器结构和跳跃连接,仍然是图像分割中实用的基线模型。然而,其瓶颈表示仍受限于有限的感受野集合,而解码器特征在传播时未明确强调最具信息量的通道和空间位置。本文提出了MSCA-UNet,一种基于U-Net的分割架构,该架构在瓶颈处结合了多尺度上下文聚合,并在解码器中引入通道-空间注意力细化。多尺度模块使用并行空洞卷积来捕获不同感受野下的上下文特征,而卷积块注意力模块(CBAM)逐步重新校准解码器特征。在相同的实验设置下,基线U-Net在留出测试集上达到96.9%的mIoU。添加多尺度上下文将mIoU提升至97.5%,而仅使用注意力则达到98.4%。结合两种机制后,mIoU达到99.1%,较基线提升了2.2个百分点。参数分析进一步表明,仅注意力的变体增加了约0.044M参数,而多尺度模块贡献了大部分额外模型容量。结果支持了多尺度上下文丰富与基于注意力的特征细化在U-Net框架内提供互补益处的观点。
英文摘要
U-Net remains a practical baseline for image segmentation because of its simple encoder-decoder structure and skip connections. However, the bottleneck representation is still dominated by a limited set of receptive fields, while decoder features are propagated without explicitly emphasizing the most informative channels and spatial locations. This paper presents MSCA-UNet, a U-Net-based segmentation architecture that combines multi-scale contextual aggregation at the bottleneck with channel-spatial attention refinement in the decoder. The multi-scale module uses parallel atrous convolutions to capture contextual features at different receptive fields, while Convolutional Block Attention Modules (CBAMs) progressively recalibrate decoder features. Under identical experimental settings, the baseline U-Net achieves 96.9% mIoU on a held-out test set. Adding multi-scale context improves mIoU to 97.5%, while attention alone reaches 98.4%. Combining both mechanisms yields 99.1% mIoU, a 2.2 percentage-point improvement over the baseline. Parameter analysis further shows that the attention-only variant adds approximately 0.044M parameters, whereas the multi-scale module contributes most of the additional model capacity. The results support the view that multi-scale context enrichment and attention-based feature refinement provide complementary benefits within a U-Net framework.
Comments5 pages, 2 figures, 1 table