发表机构
ISITCOM, University of Sousse(苏塞大学ISITCOM)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何在不破坏网络功能前提下插入残差块扩展容量。提出精确网络手术,证明相关定理,识别退化配置。通过NeuroDSL验证,如嫁接精确、门动态符合预测、成本与锥大小相关等,为网络扩展提供有效方法。
AI 中文摘要
诸如Net2Net和渐进堆叠等保持功能的网络增长技术可在不破坏模型学习功能的情况下扩展其容量,但现有公式要么容忍数值扰动,要么需要完全重建训练程序。我们形式化了精确网络手术:将残差块就地插入实时计算图,使得(i)网络功能得以保留——在显式浮点假设下逐位精确保留,并且(ii)插入的参数在插入后立即保持可训练。我们证明了门控残差块的恒等态射定理、结构局部性定理,该定理表明反应式无效引擎精确地重新计算插入点的下游锥,而不影响其他节点的值和优化器状态,以及逃逸初始化命题,该命题表明在随机初始化的分支上初始化为零的梯度阴影门α在插入时通常会接收到非零梯度。我们识别出一种退化配置——零初始化输出投影与零门相结合,这是梯度下降无法逃逸的精确鞍点。每个断言都在NeuroDSL(Julia中的反应式图引擎)的参考实现中得到了验证:嫁接在每个测试的逻辑输出上都是逐位精确的(1600个中0个不匹配);门在第一个优化器步骤中逃离零,并在第二个步骤中解锁分支梯度,正如预测的那样;退化配置在整个600步运行中梯度都恒为零;手术成本与下游锥大小的跟踪系数r = 0.9992,而嫁接加无效簿记在不同插入深度下是恒定的(约0.75毫秒);并且训练在实际过程重新启动时逐位相同地恢复。一个标记的初步附录报告了关于插入后门动态的首次单种子观察结果。
英文摘要
Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its learned function, but existing formulations either tolerate numerical perturbations or require a full rebuild of the training program. We formalize Exact Network Surgery: the in-place insertion of a residual block into a live computational graph such that (i) the network function is preserved -- bit-exactly under explicit floating-point hypotheses -- and (ii) inserted parameters remain trainable immediately after insertion. We prove an identity-morphism theorem for gated residual blocks, a structural-locality theorem showing that a reactive invalidation engine recomputes exactly the downstream cone of the insertion point, leaving every other node's value and optimizer state untouched, and an escape-from-initialization proposition showing that the Gradient Shadowing gate alpha, initialized at zero over a randomly initialized branch, receives a generically non-zero gradient at insertion time. We identify a degenerate configuration -- zero-initialized output projections combined with a zero gate -- that is an exact saddle point gradient descent cannot escape. Every claim is validated on the reference implementation in NeuroDSL, a reactive graph engine in Julia: grafting is bit-exact on every logit tested (0 mismatches out of 1600); the gate escapes zero at the first optimizer step and unlocks branch gradients at the second, exactly as predicted; the degenerate configuration exhibits gradients identically zero for the entire 600-step run; surgery cost tracks downstream cone size with r = 0.9992 while graft-plus-invalidation bookkeeping is constant (about 0.75 ms) across insertion depths; and training resumes bit-identically across a real process restart. A flagged preliminary appendix reports first single-seed observations on post-insertion gate dynamics.
CommentsCompanion paper: "Cost Accounting for Reactive Computational Graphs" (submitted concurrently)