arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MCL:元卷积层

MCL: Meta Convolution Layer

Naim Reza, Md Al Amin, Ho Yub Jung

arXiv 2610.11117首次发表:更新:

发表机构

Chosun University(朝鲜大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出元卷积层(MCL),将卷积核建模为输入条件函数,解耦表征能力与混合核数量,在多个图像分类任务中提升了ResNet等模型的准确率,性能优于现有动态卷积方法。

AI 中文摘要

动态卷积通过使卷积核适配输入内容来增强卷积神经网络(CNN),但它将有效卷积核表示为少量基卷积核的线性混合,这限制了表达能力,且随着混合规模增大优化变得复杂。本研究从函数视角重新审视动态卷积,提出元卷积层(Meta Convolution Layer, MCL),该层将卷积核直接建模为通过高阶多项式展开实现的输入条件函数W(x)。利用受深度多项式网络启发的嵌套残差块,MCL实现了结构化多项式元网络,生成单个输入自适应卷积核,从而将表征能力与显式混合卷积核数量解耦,缓解了训练不稳定性。MCL是可与标准卷积结合的即插即用模块,可无缝集成到CNN和Transformer主干网络中。实验评估显示,在ImageNet数据集上,添加MCL可将ResNet-18、ResNet-50和ResNet-101的Top-1准确率分别提升6.61%、3.42%和3.05%;此外,该方法显著提高了ResNet和Wide-ResNet变体在CIFAR-10和CIFAR-100数据集上的准确率,在使用Swin和ViT主干网络的细粒度视觉分类任务中,其性能也优于现有方法。这些结果表明,高阶多项式卷积核生成是基于线性混合的动态卷积的强大且可扩展的替代方案。

英文摘要

Dynamic convolution enhances convolutional neural networks (CNNs) by adapting kernels to input content, but it expresses the effective kernel as a linear mixture of a small number of basis kernels, which limits expressivity and complicates optimization as the mixture size grows. In this work, we revisit dynamic convolution from a functional perspective and propose the Meta Convolution Layer (MCL), which directly models the convolutional kernel as an input-conditioned function W(x) realized via a high-order polynomial expansion. Leveraging nested residual blocks inspired by deep polynomial networks, MCL implements a structured polynomial meta-network that generates a single input-adaptive kernel, thereby decoupling representational power from the explicit number of mixture kernels and alleviating training instability. MCL is a plug-in addition with standard convolutions and can be seamlessly integrated into both CNN and transformer backbones. Experimental evaluation shows that adding MCL improves the Top-1 accuracy of Resnet- 18, Resnet-50 and ResNet-101 by 6.61%, 3.42% and 3.05% on the ImageNet dataset. Moreover, the proposed method significantly boosts the accuracy of Resnet and Wide-Resnet variants on CIFAR-10 and CIFAR-100 datasets. Additionally, the proposed method outperforms previous methods on fine-grained visual classification tasks using Swin and ViT backbones. These results demonstrate that high-order polynomial kernel generation is a powerful and scalable alternative to linear mixture based dynamic convolution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑