全连接层的分层解耦稳定结构化稀疏化方法
Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers
浏览论文内容
中文总结 AI 辅助
提出一种逐层解耦的结构化稀疏化方法,通过提取两层子网络并施加组惩罚,实现更稳健的剪枝,具有更宽的正则化范围和更低的过度剪枝率。
中文摘要 AI 辅助
我们提出了一种解耦的、逐层处理的方法,用于对预训练神经网络的全连接层进行结构化稀疏化。与对所有层进行联合惩罚不同,我们的方法提取浅层两层子网络,对内层权重进行归一化,并对每个块的外层权重矩阵施加结构化组惩罚,按顺序处理各层以剪除神经元并减小每层的宽度。我们证明了在任意正齐次激活函数下,受约束的解耦目标在最优性上等价于对内层和外层权重的特定联合惩罚,因此具有简洁的投影和近端公式。我们的核心发现是,这种解耦重构比耦合方法更稳健。在数值实验中,与测试的联合基线相比,它提供了更宽的正则化强度可用范围,并降低了灾难性过度剪枝的发生率,同时保持了相当的精度。我们在受控的分类和稀疏恢复研究中确立了这些性质,并在高维PINN压力测试以及OPT-1.3B的前馈层中考察了其适用范围。
英文摘要
We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.
发表机构
- The University of Scranton(斯克兰顿大学)
- University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
机构由 AI 辅助整理,请以论文原文为准。