持续学习的序列代价
The Sequential Price of Continual Learning
浏览论文内容
中文总结 AI 辅助
本研究在过参数化线性回归模型中证明持续学习的序列更新产生额外代价,其平稳极限可精确分解,并分析EWC的缓解作用,在Jester数据集上验证理论。
中文摘要 AI 辅助
序列任务更新是持续学习的基础,但其对近期任务的偏向可能造成持久的性能代价。我们在一个具有独立同分布任务采样的过参数化线性回归模型中研究这一代价。我们证明了分布层面的遗忘和总体损失收敛到相同的平稳极限。这一共同极限精确地分离为联合训练渐近达到的内在损失和额外的序列代价,并且在更同质的任务几何中,这两项重合,使得总损失是联合训练的两倍。我们进一步分析了在一般任务曲率下固定强度的弹性权重巩固(EWC),并刻画了其在每个正则化强度下的平稳序列代价。在强正则化下,代价随EWC强度呈反比衰减,而收敛到平稳性的速度在同一尺度上减慢。在Jester笑话评分数据集上,该理论精确量化了由自然冲突的用户偏好产生的序列代价以及EWC对其的减少。
英文摘要
Sequential task updates are fundamental to continual learning, but their recency bias can impose a lasting performance cost. We study this cost in an overparameterized linear-regression model with i.i.d. task sampling. We prove that distribution-level forgetting and population loss converge to the same stationary limit. We quantify the additional loss incurred by sequential exact fitting, or the sequential price. In more homogeneous task geometries, it equals the intrinsic loss asymptotically attained by joint training, making the total loss twice as large. We further analyze fixed-strength elastic weight consolidation (EWC) under general task curvatures and characterize its stationary sequential price at every regularization strength. Under strong regularization, the price decays inversely with EWC strength while the mean-square coupling horizon grows proportionally. Experiments on Jester and Rotated MNIST support the predicted sequential price and its reduction by EWC, with quantitative agreement on real-world tasks satisfying the theory's assumptions and qualitative agreement under nonlinear finite-step training.
发表机构
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。