层很重要:为何持续学习正则化应是层自适应的
Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive
浏览论文内容
中文总结 AI 辅助
该研究针对持续学习正则化的层自适应问题,提出保护早期层、允许深层移动的准则,应用于EWC和SLCA后提升了平均性能与遗忘指标。
中文摘要 AI 辅助
EWC等持续学习正则化方法通过按参数重要性(通常为对角Fisher值)惩罚与前一任务参数的变化来对抗遗忘。按参数调整看似比按层调整更灵活,但每层的对角Fisher值是其实际曲率的弱摘要,缺失控制遗忘的最大特征值信息。对抗性位翻转攻击和Hessian谱研究表明,这种缺失的每层敏感性在神经网络中跨越多个数量级。在块对角Hessian假设(即EWC现有对角假设的层级类似假设)下,我们证明三点:遗忘可分解为按每层最大Hessian特征值加权的层项之和;对角Fisher权重无法恢复该特征值;例如,两个Fisher平均值相同的层,其最大特征值差异可达层宽度的倍数。在相同遗忘水平下,均匀正则化会使新任务性能损失,损失量与层条件数成比例。我们的理论分析得出简单准则:强保护早期层,允许深层移动。将该准则应用于EWC和SLCA,结果显示平均性能和遗忘指标均有明显提升。
英文摘要
Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values. Per-parameter looks more flexible than per-layer, but each layer's diagonal Fisher is a weak summary of its actual curvature, missing the top-eigenvalue information that controls forgetting. Adversarial bit-flip attacks and Hessian-spectrum studies show that this missing per-layer sensitivity spans orders of magnitude in neural networks. Under a block-diagonal Hessian assumption, the layer-level analogue of EWC's existing diagonal assumption, we prove three things. Forgetting decomposes as a sum of per-layer terms weighted by each layer's top Hessian eigenvalue. Diagonal-Fisher weights cannot recover this eigenvalue. For instance, two layers with identical Fisher averages can have top eigenvalues differing by a factor as large as the layer width. For the same level of forgetting, uniform regularization loses new-task performance by an amount scaling with the layer condition number. Our theoretical analysis leads to a simple recipe: protect early layers strongly, let deeper layers move. We apply this recipe to EWC and SLCA and show clear improvements in average performance and forgetting metrics.
发表机构
- German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
- RPTU University Kaiserslautern-Landau(凯撒斯劳滕-兰道工业大学(RPTU))
机构由 AI 辅助整理,请以论文原文为准。