训练过程中简化神经网络
Simplifying Neural Networks During Training
- University of Turin(都灵大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究结合神经坍缩与隧道效应,提出受神经坍缩启发的训练框架,通过逆费舍尔准则确定简化节点,将后续层替换为轻量级分类头,在MLP、VGG等模型上实现大幅参数缩减且保持准确率。
AI中文摘要:
理解和利用过参数化深度神经网络的训练动态仍是现代机器学习的核心挑战。关于神经坍缩(Neural Collapse, NC)的最新证据表明,类别表示和分类器呈现高度结构化的几何特性;而隧道效应(Tunnel Effect)则指出,仅有部分层对特征提取至关重要。我们结合这两个视角,提出一种受神经坍缩启发的训练框架,用于在训练过程中简化深度网络。该方法通过逆费舍尔准则(Inverse Fisher Criterion)监测表示动态,该准则是变异性坍缩行为的稳定且高效的替代指标,用于识别特征提取与分类之间的分界点,以及简化可行的训练阶段。随后,我们将后续层替换为轻量级分类头,并继续训练简化后的模型。在MLP、VGG和ResNet架构的图像分类基准上的实验表明,所提方法在实现大幅参数缩减的同时,保持了与完整模型相当的准确率。可在此URL获取复现实验的代码。
英文摘要:
Understanding and exploiting the training dynamics of overparameterized deep neural networks remains a central challenge in modern machine learning. Recent evidence on Neural Collapse (NC) shows that class representations and classifiers exhibit highly structured geometry, while the Tunnel Effect suggests that only a subset of layers is essential for feature extraction. We combine these two perspectives and propose an NC-inspired training framework for simplifying deep networks during training. Our method monitors representation dynamics through the Inverse Fisher Criterion, a stable and efficient proxy for the variability collapse behavior, to identify both the split point between feature extraction and classification and the training stage at which simplification becomes viable. We then replace the trailing layers with a lightweight classification head and continue training the reduced model. Experiments on image-classification benchmarks across MLP, VGG, and ResNet architectures show that the proposed method achieves substantial parameter reductions while maintaining accuracy comparable to that of the full model. Code to reproduce the experiments can be found at: https://github.com/LorenzoSciandra/NNS.