arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10387cs.CV

增强型可变形卷积:中心不变偏移与边缘感知掩码

Enhanced Deformable Convolution with Center-invariant Offset and Edge-aware Mask

  • Beihang University(北京航空航天大学)
  • Shanghai Academy of Spaceflight Technology(上海航天技术研究院)
  • Shanghai Radio Equipment Research Institute(上海无线电设备研究所)
  • Cardiff University(卡迪夫大学)
  • Shenzhen University(深圳大学)
  • Sun Yat-sen University(中山大学)
  • Lancaster University(兰卡斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

Yixiao Li, Xiaoyuan Yang, Jin Jiang, Minghao Zou, Guanghui Yue, Baoquan Zhao, Jun Liu, Wei Zhou

AI总结:

本文提出增强型可变形卷积(EDC),通过中心不变偏移模块和边缘感知掩码模块,改进语义分割中的可变形卷积,提升空间适应性和目标聚焦,优于现有变体。

AI中文摘要:

可变形卷积网络近年来在许多计算机视觉任务中变得流行,尤其是在语义分割中,因为它们在动态空间建模方面具有卓越的能力。然而,由于密集的可变形偏移和缺乏更长范围的依赖,它们无法完全采用适当且精确的变形来进行特征表示。为了解决这些问题,本文提出了用于语义分割的增强型可变形卷积网络(EDCN)。具体而言,在解码器中采用了一种新颖的增强型可变形卷积(EDC),它集成了中心不变偏移模块(COM)和边缘感知掩码模块(EMM)。COM使用更大的卷积核并消除卷积核中心的变形,从更丰富的空间信息中获得更符合目标的偏移。同时,EMM通过Sobel边缘检测获取图像内容的重要性,然后根据内容重要性选择性地应用变形,最小化与相对不重要信息相关的不必要变形,从而避免来自信息量较少区域的影响。实验表明,EDC在主流分割数据集和不同解码器设置下,优于最先进的可变形卷积变体,包括Deformable ConvNets V1-V4和Entire Deformable ConvNets。此外,消融研究证实了每个组件的有效性。另外,可视化结果表明EDC增强了空间适应性和目标聚焦。我们进一步分析了EDC在图像分类基准上扩展到更大卷积核的可行性。代码将公开发布。

英文摘要:

Deformable convolution networks have recently become popular for many computer vision tasks, especially for semantic segmentation, because of their exceptional capabilities in dynamic spatial modeling. However, due to the dense deformable offsets and the lack of longer-range dependencies, they can not fully adopt proper and precise deformations for feature representations. To tackle the issues, in this paper, we propose Enhanced Deformable ConvNets (EDCN) for semantic segmentation. Specifically, a novel Enhanced Deformable Convolution (EDC) is exploited in the decoder, which integrates the Center-invariant Offset Module (COM) and Edge-aware Mask Module (EMM). The COM employs larger kernels and eliminates deformations at the kernel center, obtaining offsets that are more in line with the target from richer spatial information. Concurrently, the EMM obtains the significance of image content via Sobel edge detection, then selectively applies deformations based on the content significance, minimizing unnecessary deformations associated with relatively less important information, thereby avoiding impact from less informative regions. Experiments show that EDC outperforms state-of-the-art deformable convolution variants, including Deformable ConvNets V1-V4 and Entire Deformable ConvNets, across mainstream segmentation datasets with various decoder settings. Moreover, ablation studies confirm the effectiveness of each component. In addition, visualizations illustrate that EDC enhances spatial adaptation and target focus. We further analyze the extendibility of EDC to larger kernels on the image classification benchmark. Code will be publicly released.

补充信息

↑