arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31630cs.LGcs.NEstat.ML

静默自由度中的重放:无离线阶段的持续学习

Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase

Yanhai Zhang, Jie Zhang, Chi Xu

AI总结:

提出一种无需离线阶段、基于局部规则的持续学习重放方法,通过隔离和不应期轮换在静默自由度中巩固记忆,在分割MNIST上达到91.6%的准确率,优于多种现有方法。

AI中文摘要:

基于重放的持续学习几乎总是在专门的离线阶段进行巩固,或者通过将重放样本与输入流交错进行,而大脑在清醒时也通过局部睡眠(即单个回路短暂的使用依赖性关闭期)进行巩固。我们探究一个由局部、生物约束规则训练的网络是否可以在完全没有离线阶段的情况下进行巩固。一种隔离规则将重放更新限制在当前输入下k-胜者全取(k-winner-take-all)动态中不可见的隐藏突触上,优化器状态仅在掩码内推进;一种不应期轮换规则使刚发放过的单元在下一轮竞争中退出,从而扩大可巩固集合;一种稳态压力和相对新颖性门控决定重放突发何时触发以及轮换何时运行。这颠倒了无干扰持续学习的通常方向:当前输入上的隐藏计算保持不变(在已证明的通道上完全不变,在其他地方每次更新中除0.3%的清醒样本外均不变),而过去的记忆被写入当前批次未使用的自由度中。在类增量分割MNIST(split-MNIST)上,系统在没有离线阶段的情况下达到91.6±0.3%,在保留的两个分割上达到或超过最佳离线夜间调度,与DER++持平,并高于经验重放、ER-ACE、A-GEM和未掩码的局部重放;在单次遍历中,它领先于DER++(91.8%对90.1%),而夜间调度降至76.9%。优势在小缓冲区时最大,在大缓冲区时让位于反向传播参考方法;在分割CIFAR-10上,系统领先于离线排练和经验重放,但落后于ER-ACE和DER++。轮换承担了大部分增益;隔离增加了不变性保证。该机制不依赖于局部规则:在相同调度下,具有k-WTA隐藏层的反向传播网络也能从轮换中获益,而隔离在其之上再次免费获得。

英文摘要:

Replay-based continual learning rehearses past data either in an offline phase, during which the agent stops acting, or interleaved with the live stream, where it perturbs the computation serving the current input. An agent that learns in deployment can afford neither. Motivated by local sleep, the use-dependent off periods of individual cortical circuits in awake animals, we show that replay can instead be written into the degrees of freedom the current input leaves unused. In a network with k-winner-take-all hidden layers, confining replay updates to synapses whose presynaptic unit is silent or whose postsynaptic unit is inactive leaves the hidden computation on the current batch invariant: exactly so for silent and suppressed units, and for all but 0.3% of samples in practice. A refractory rule under which units that have just fired sit out the next competition doubles the width of this channel and carries most of the accuracy. On class-incremental split-MNIST the resulting learner, with no offline phase, matches or exceeds the best offline rehearsal schedule and outperforms experience replay, ER-ACE and unmasked interleaved replay, each re-tuned under the same micro-batch schedule. Against DER++ the comparison splits by protocol: with five epochs per task DER++ leads by 1.5 points once it runs under that schedule, and in a single pass over the stream, the regime closest to the agent deployment setting, the system leads it by 1.6 while offline rehearsal falls 15 points behind. On split CIFAR-10 it again leads offline rehearsal and experience replay, but trails ER-ACE and DER++ by two to three points. The construction is not tied to the local learner: on a backprop network with k-winner-take-all hidden layers under the same schedule, refractory rotation adds half a point in five epochs and three in a single pass, and isolation again costs nothing on top of it.

补充信息

↑