发表机构
Votee AI; Beever AI(Votee人工智能公司; Beever人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出门零增长这一用于持续学习的函数保持算子,通过零初始化门添加残差块,在横截性条件下实现秩分离。实验表明其能控制函数漂移和雅可比矩阵泄漏,在变压器模型上旧域遗忘近零,优于非FP控制,还涵盖多种结构确立其为规范实例。
AI 中文摘要
我们引入了门零增长,这是一种用于持续学习的函数保持(FP)算子,它通过零初始化门添加新的残差块。在横截性条件下,门零增长在函数雅可比矩阵中诱导出秩分离:旧方向不变,新权重方向在增长点处恰好平坦,新门方向是新函数变化的唯一一阶来源。随着门在持续学习中打开,函数漂移为$O(\|\boldsymbol{\alpha}\|^2)$,雅可比矩阵泄漏为$O(\|\boldsymbol{\alpha}\|_\infty)$,从而实现与FP轨迹的可控偏离。在从WikiText - 103改编到BookCorpus的300M→857M变压器上,门零增长在精确保持(隔离)和联合前沿(不冻结)操作点下都达到了接近零的旧域遗忘($\Delta_A < 0.1$),而非FP控制($G_{\text{stack}}$)在相同方法下遗忘程度大一个数量级。相同的几何分析涵盖了LoRA、ReZero和零初始化适配器结构,将门零增长确立为共享局部几何的规范实例,该几何控制着持续学习中的安全容量激活。
英文摘要
Model growth adds parameters to a trained checkpoint and continues training, aiming to increase capacity while reusing the existing model instead of training from scratch. But does growing a trained model actually create new functional capacity that survives subsequent training? Parameter count alone cannot answer this question, and ``capacity'' is itself ambiguous; it can refer to parameter count, effective dimensionality, or the ability to represent new functions. We measure it directly, as the number of new, functionally independent directions a growth step adds, formalized as the rank increment of the model's functional Jacobian and estimated matrix-free at scale. For zero-gated function-preserving growth, this increment is exactly $K$ under a transversality condition, i.e. $\operatorname{rank}(J_{\mathrm{grown}})=\operatorname{rank}(J_{\mathrm{old}})+K$. Continually training a Transformer grown from $253\mathrm{M}$ to $857\mathrm{M}$ parameters across three domains, we find these added directions persist, with all $K=36$ remaining independent after continual learning, even when the model catastrophically forgets (perplexity $28\to481$) or we deliberately destroy the old function ($25.9\to127$). Geometric capacity and retained function are therefore decoupled. The functional capacity introduced by growth persists, while the previously learned function is not retained. A second, geometrically distinct mechanism, Net2Net, reaches the same $K$-dimensional capacity from directions that are dormant at initialization ($0 \to K$) and then retains it, so the decoupling is not specific to the zero-gate parameterization. The implication is precise. Catastrophic forgetting need not reflect a loss of functional capacity, and preserving dimensionality alone is insufficient to preserve the learned function.
Comments18 pages, 1 figures