arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于大核卷积神经网络中移动高效逐点卷积的组共享低秩近似

Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs

Hao Luo, Yiting Yang, Wenyi Zhao, Man Jiang, Zhijun Lin, Ghulam Mohiuddin, Ting Jiang, Kunming Luo, Zihao Zhang, Qingsen Yan, Guoqing Wang, Wei Dong, Peng Wang

arXiv 2608.26069首次发表:更新:

发表机构

Northwestern Polytechnical University; Nanchang University; China Mobile Chengdu Institute of Research and Development; Hong Kong University of Science and Technology; Institute of AI for Industries, Chinese Academy of Sciences; University of Electronic Science and Technology of China(西北工业大学; 南昌大学; 中国移动成都研究院; 香港科技大学; 中国科学院人工智能产业研究院; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大核CNN中逐点卷积占比过高导致边缘部署瓶颈的问题,提出通道组共享低秩近似方法,在保持性能的同时降低存储成本,实现大核CNN的边缘高效部署。

AI 中文摘要

大核卷积神经网络(CNNs)通过显著扩大感受野在视觉任务中展现出卓越性能,但其参数的二次增长严重阻碍了存储高效的边缘部署。尽管现有高效架构采用了参数高效的深度可分离卷积主干,利用低秩近似和权重共享等技术压缩深度卷积,但我们发现了一个关键疏漏:逐点卷积占据了参数总量的绝大部分(如RepLKNet-31B模型中占比超过87%),是资源受限边缘设备(如内存为4-12GB随机存取存储器(RAM)的智能手机)上的主要部署瓶颈,这导致资源有限设备上出现难以承受的存储成本和严重的内存加载约束。为解决该问题,我们提出了通道组共享(CGS)低秩近似,这是一种基于奇异值分解(SVD)的新型参数共享策略。CGS构建了与SVD分解同构的结构化低秩范式,包含层内跨通道组共享的(高参数成本)下/上投影矩阵,以及通道组特有的(低参数成本)可扩展对角矩阵。这种组共享设计实现了显著的参数减少。大量实验表明,采用CGS增强的大核CNNs(RepLKNet、ConvNeXt、SLaK)在保持竞争力性能与大幅降低存储成本之间取得了经验上的良好平衡。至关重要的是,通过缓解存储约束、减少加载时的内存带宽压力并最小化模型加载延迟,CGS使预训练大核CNN模型能在边缘设备上可行部署,从而弥合了高性能视觉模型与实际边缘部署之间的差距。

英文摘要

Large-kernel Convolutional Neural Networks (CNNs) deliver remarkable performance in vision tasks by significantly expanding receptive fields, yet their quadratic parameter growth critically impedes storage-efficient edge deployment. While existing efficient architectures adopt parameter-efficient depthwise separable convolution backbones that leverage techniques like low-rank approximation and weight sharing to compress depthwise convolutions, we identify a critical oversight: pointwise convolutions dominate parameter volume (>87% in models like RepLKNet-31B) and constitute the primary deployment bottleneck on resource-constrained edge devices. This results in prohibitive storage costs and severe memory-loading constraints on resource-limited devices (e.g., smartphones with 4-12 GB Random Access Memory (RAM)). To overcome this, we propose Channel Group-Shared (CGS) low-rank approximation, a novel Singular Value Decomposition (SVD)-based parameter-sharing strategy. CGS constructs a structured low-rank paradigm isomorphic to SVD decomposition, comprising shared (high-parameter-cost) down/up-projection matrices across channel groups within a layer and channel-group-specific (low-parameter-cost) scalable diagonal matrices. This group-sharing design achieves significant parameter reduction. Extensive experiments demonstrate that large-kernel CNNs (RepLKNet, ConvNeXt, SLaK) enhanced with CGS strike an empirically favorable balance between competitive performance and substantially reduced storage costs. Crucially, by alleviating storage constraints, reducing memory bandwidth pressure during loading, and minimizing model loading latency, CGS enables the feasible deployment of pre-trained large-kernel CNN models on edge devices, thereby bridging the gap between high-performance vision models and practical edge deployment.

Comments17 pages, 10 figures, accepted by MobiCom2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑