arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14576cs.LGstat.ML

深度残差架构的尖锐稳定性阈值与认证

Certifying Residual Architectures from Their Primitives: A Sharp Stability Threshold

Hyemin Gu, Michael Tyrrell, Tuhin Sahai, Markos A. Katsoulakis

首次发表
浏览论文内容

中文总结 AI 辅助

该研究为深度残差架构提出次线性增长原理,通过经典ODE理论和最优控制分析确定\(q \leq 1\)为稳定训练充要条件,阐明架构布局与稳定关系,构建函数空间并给出认证算法,实验证实相关变体训练稳定。

中文摘要 AI 辅助

我们为深度残差架构提出了次线性增长原理——每个残差块速度场的输入幅度指数上的尖锐稳定性阈值:\(\|v(x, t)\| \leq c\,\|x\|^q + b\),\(q \in [0, 1]\)。通过两个独立论证确定阈值\(q = 1\)。经典常微分方程理论在\(q \leq 1\)时给出\([0, T]\)上的全局前向流,\(q > 1\)时速度场发散。最优控制分析通过哈密顿 - 雅可比 - 贝尔曼方程将其细化为选择陈述:训练最优在可允许类边界上是开关控制的,所以\(q > 1\)时最优解爆炸,\(q \leq 1\)时最优解通过构造是安全的。指数准则\(q \leq 1\)是稳定训练的充要条件。它阐明了确保训练和推理稳定性的架构布局,解释了例如层归一化的稳定作用。次线性增长速度场构成了正向动力学、伴随灵敏度和架构组合都能得到良好控制的正确函数空间。在构建残差块的五种操作下的输入幅度指数算法能够在架构原语层面有效认证\(q_k \leq 1\),取代在寻找稳定神经架构设计中的临时试错。无参数修改将超临界曼巴块从\(q = 5\)降至\(q = 1\)且无需层归一化,证明了这一点。在曼巴和PatchTST上的实验证实\(q \leq 1\)变体训练稳定:准则是输入幅度指数,而非归一化层的存在。

英文摘要

Whether a deep residual architecture trains stably is usually determined by training it, which is expensive and answers the question only for the architecture, depth, and floating-point format tested. In practice, stability is secured by heuristics for where to place normalization and which type to use, supported by experiments but lacking a common principle. We show that stability can instead be certified before training, directly from the architectural primitives of a residual block, as an explicit function of depth and floating-point format. The certificate has two elements: (a) a power-law growth bound, $|v(x)|\le c|x|^q+b$, whose exponent $q$ is computed from the block's primitives by an arithmetic of exponents (forward); and (b) gradient bounds from a Lipschitz condition on the block over reachable states (backward). The certificate determines when the state can reach the largest finite floating-point value $M$: for $q\le1$, overflow requires at least order $\log M$ layers, whereas for $q>1$ it can occur within order $\log_q\log M$ layers; the threshold $q=1$ is sharp. Every normalization and bounded activation sets $q=0$, and the arithmetic identifies minimal modifications that bring a block from $q>1$ to $q=1$ without normalization. Empirically, in forecasting, only $q>1$ blocks blow up, as predicted by the forward certificate. In GPT-2 on OpenWebText, $q=1$ blocks without normalization also diverge: the forward certificate holds, while the backward one fails in the attention block with the largest query-key coefficient. Relaxing normalization from $q=0$ to $q=1$ improves out-of-distribution generalization in operator learning and accuracy in time-series forecasting. As deep models become more complex and costly to train, our results provide a way to assess the trade-off between stability and representational flexibility directly from block primitives, before training.

发表机构

  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • SRI International(SRI国际公司)

机构由 AI 辅助整理,请以论文原文为准。

↑