arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MSCA-UNet:用于图像分割的多尺度上下文与注意力U-Net

MSCA-UNet: Multi-Scale Context and Attention U-Net for Image Segmentation

Sheng-Wei Chan

arXiv 2609.06356首次发表:更新:

发表机构

Tamkang University(淡江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MSCA-UNet,在U-Net瓶颈处结合多尺度空洞卷积上下文聚合与解码器CBAM注意力细化,将mIoU从基线96.9%提升至99.1%,验证了两者的互补增益。

AI 中文摘要

U-Net因其简单的编码器-解码器结构和跳跃连接,仍然是图像分割中实用的基线模型。然而,其瓶颈表示仍受限于有限的感受野集合,而解码器特征在传播时未明确强调最具信息量的通道和空间位置。本文提出了MSCA-UNet,一种基于U-Net的分割架构,该架构在瓶颈处结合了多尺度上下文聚合,并在解码器中引入通道-空间注意力细化。多尺度模块使用并行空洞卷积来捕获不同感受野下的上下文特征,而卷积块注意力模块(CBAM)逐步重新校准解码器特征。在相同的实验设置下,基线U-Net在留出测试集上达到96.9%的mIoU。添加多尺度上下文将mIoU提升至97.5%,而仅使用注意力则达到98.4%。结合两种机制后,mIoU达到99.1%,较基线提升了2.2个百分点。参数分析进一步表明,仅注意力的变体增加了约0.044M参数,而多尺度模块贡献了大部分额外模型容量。结果支持了多尺度上下文丰富与基于注意力的特征细化在U-Net框架内提供互补益处的观点。

英文摘要

U-Net remains a practical baseline for image segmentation because of its simple encoder-decoder structure and skip connections. However, the bottleneck representation is still dominated by a limited set of receptive fields, while decoder features are propagated without explicitly emphasizing the most informative channels and spatial locations. This paper presents MSCA-UNet, a U-Net-based segmentation architecture that combines multi-scale contextual aggregation at the bottleneck with channel-spatial attention refinement in the decoder. The multi-scale module uses parallel atrous convolutions to capture contextual features at different receptive fields, while Convolutional Block Attention Modules (CBAMs) progressively recalibrate decoder features. Under identical experimental settings, the baseline U-Net achieves 96.9% mIoU on a held-out test set. Adding multi-scale context improves mIoU to 97.5%, while attention alone reaches 98.4%. Combining both mechanisms yields 99.1% mIoU, a 2.2 percentage-point improvement over the baseline. Parameter analysis further shows that the attention-only variant adds approximately 0.044M parameters, whereas the multi-scale module contributes most of the additional model capacity. The results support the view that multi-scale context enrichment and attention-based feature refinement provide complementary benefits within a U-Net framework.

Comments5 pages, 2 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑