arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当简洁取胜:面向轻量级语义分割的瓶颈感知上下文建模

When Simplicity Wins: Bottleneck-Aware Context Modeling for Lightweight Semantic Segmentation

Mian Muhammad Naeem Abid, Nancy Mehta, Zongwei Wu, Radu Timofte

arXiv 2608.18979首次发表:更新:

发表机构

University of Würzburg(维尔茨堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对轻量级语义分割的瓶颈阶段被忽视问题,提出含三个互补组件的SiConMo框架,在四个数据集上实现最优精度-效率权衡,凸显简洁性设计原则的有效性。

AI 中文摘要

语义分割需要在精度、效率和可扩展性之间进行谨慎权衡,这对于高分辨率图像来说仍然难以实现。卷积网络能有效建模局部模式,但难以处理长程依赖;而视觉Transformer虽能捕获全局上下文,但计算成本很高。尽管近期研究大多聚焦于编码器设计,但作为上下文聚合与信息流核心的瓶颈阶段却被相对忽视。我们提出SiConMo,这是一个轻量且有效的框架,包含两个变体:仅RGB模型(SiConMo)和增强GME的变体(SiConMo$_\text{†}$)。我们证明简洁性源于关键设计原则:在极低计算预算下,瓶颈是整合局部与全局上下文的最高效阶段。SiConMo整合了三个互补组件:用于分层多尺度表示的Token金字塔提取模块、用于瓶颈感知上下文建模的Transformer分支深度卷积块,以及在保留空间结构同时增强语义一致性的特征融合模块。在ADE20K、PASCAL Context、Cityscapes和COCO-Stuff四个数据集上的大量实验表明,SiConMo在轻量级语义分割模型中实现了最先进的精度-效率权衡,凸显简洁性是一种强大的设计原则。

英文摘要

Semantic segmentation demands a careful balance between accuracy, efficiency, and scalability, which remains difficult to achieve for high-resolution imagery. Convolutional networks effectively model local patterns but struggle with long-range dependencies, whereas Vision Transformers capture global context at a high computational cost. While recent work largely focuses on encoder design, the bottleneck stage, central to contextual aggregation and information flow, has been relatively overlooked. We propose SiConMo, a lightweight yet effective framework, implemented in two variants: an RGB-only model (SiConMo) and a GME-enhanced variant (SiConMo$_\dagger$). We show that simplicity arises from a key design principle: at very low computational budgets, the bottleneck is the most efficient stage to integrate local and global context. SiConMo integrates three complementary components: a Token Pyramid Extraction Module for hierarchical multi-scale representation, a Transformer-Branched Depthwise Convolution block for bottleneck-aware context modeling, and a Feature Merging Module that preserves spatial structure while enhancing semantic consistency. Extensive experiments on ADE20K, PASCAL Context, Cityscapes, and COCO-Stuff demonstrate that SiConMo achieves a state-of-the-art accuracy-efficiency trade-off among lightweight semantic segmentation models, highlighting simplicity as a powerful design principle.

CommentsAccepted at IEEE ICIP 2026; ranked among the Top 3%

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑