AI 中文总结
该研究提出一阶准则量化残差神经网络的深度充分性,证明残差非退化是深度有一阶价值的充要条件,验证激活梯度幅度可作为深度剩余一阶价值的保守诊断指标。
AI 中文摘要
如何判断一个已训练的神经网络是否已足够深?我们在固定的函数保持型残差增长协议下研究该问题,该协议指定了插入位置、残差族、零输出初始化以及零状态一阶更新。我们将一阶残差深度饱和定义为:不存在任何可允许的插入能产生严格局部下降。我们证明残差非退化是必要且充分条件:当条件激活梯度在至少一个可允许的残差切空间上具有非零投影时,额外深度具有一阶价值。该边界与可下降兼容的零状态更新共享,且在保持该切空间的正则局部重参数化下不变。在残差信号可实现性下,原始激活梯度消失恰好可证明饱和。在ResNets、GPT-2风格模型及持续预训练的Pythia检查点中,最大激活梯度范数随深度向低信号区域减小。函数保持型增长也能达到与从头训练相当的收敛性能。这些结果支持将激活梯度幅度作为残差深度剩余经验一阶价值的保守诊断指标。
英文摘要
How can we determine whether a trained neural network is already deep enough? We study this under a fixed function-preserving residual-growth protocol specifying insertion locations, residual families, zero-output initializations, and zero-state first-order updates. We define first-order residual depth saturation as the absence of a strict local decrease from every admissible insertion. We prove residual non-degeneracy is necessary and sufficient: additional depth has first-order value exactly when conditional activation gradients have a nonzero projection onto at least one admissible residual tangent space. This boundary is shared by descent-compatible zero-state updates and invariant under regular local reparameterizations preserving that tangent space. Under residual-signal realizability, raw activation-gradient vanishing exactly certifies saturation. Across ResNets, GPT-2-style models, and continued-pretrained Pythia checkpoints, the maximum activation-gradient norm decreases toward a low-signal regime with depth. Function-preserving growth also achieves converged performance competitive with training from scratch. These results support activation-gradient magnitude as a conservative diagnostic of the remaining empirical first-order value of residual depth.