AI 中文总结
针对神经网络过参数化带来的计算与内存开销问题,提出受反赫布规则启发的DADP剪枝方法,通过全局阈值动态分配稀疏度,在多种架构上性能优于或相当现有方法,且能保留特征多样性。
AI 中文摘要
现代神经网络存在严重的过参数化问题,这种冗余会在训练和推理阶段带来大量的计算与内存开销。现有的剪枝方法依赖事后的幅值阈值或静态初始化启发式规则,因此往往需要手动设置每层的稀疏度目标或进行昂贵的重训练循环。我们提出动态活动依赖剪枝(Dynamic Activity-Dependent Pruning, DADP),这是一种受生物学启发的结构化可塑性机制。在训练过程中,DADP通过突触前激活值与突触后误差梯度的累积乘积来衡量连接的重要性。DADP使用单一全局阈值而非固定的层预算,可在网络深度上动态分配稀疏度,同时自然诱导神经元级和通道级剪枝。在MLP、VGG-16、ResNet-18、BiLSTM-CRF和MiniBERT架构上,DADP的性能与Magnitude、SNIP和RigL相当或更优;在ResNet-18上,当稀疏度达到99%时,DADP保留了73.67%的准确率(密集基线为76.06%)。最后,基于矩阵的香农熵和有效秩测量结果证实,DADP在极高稀疏度下能保留潜在特征多样性,不会出现表示崩溃。
英文摘要
Modern neural networks are heavily over-parameterized. This redundancy incurs substantial compute and memory overhead during training and inference. Existing pruning methods rely on post-hoc magnitude thresholds or static initialization heuristics. Consequently, they often require manual per-layer sparsity targets or expensive retraining cycles. We propose Dynamic Activity-Dependent Pruning (DADP), a biologically inspired structural plasticity mechanism. During training, DADP measures connection importance via the accumulated product of pre-synaptic activations and post-synaptic error gradients. Using a single global threshold instead of fixed layer budgets, DADP dynamically allocates sparsity across network depth while naturally inducing neuron- and channel-level pruning. Across MLP, VGG-16, ResNet-18, BiLSTM-CRF, and MiniBERT architectures, DADP matches or outperforms Magnitude, SNIP and RigL, retaining 73.67% accuracy (dense baseline: 76.06%) at 99% sparsity on ResNet-18. Finally, matrix-based Shannon entropy and effective rank measurements confirm that DADP preserves latent feature diversity at extreme sparsities without representation collapse.
Comments24 Pages, 8 figures