arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无重放持续学习中的噪声最优:隔离、机制与适用范围

A Noise Optimum in Rehearsal-Free Continual Learning: Isolation, Mechanism, and Scope

Gunner Levi Howe

arXiv 2609.20162首次发表:更新:

AI 中文总结

本文通过模拟隔离了无重放持续学习中保持能力与噪声间的倒U型最优,发现其源于向整合权重的相干恢复及锚定增益与噪声方差耦合,并界定其需共享任务结构,为噪声注入规则提供机制解释。

AI 中文摘要

向整合规则中注入随机噪声可以提高网络对早期任务的保持能力,直至达到最优水平,然后降低——即保持能力与噪声之间呈倒U型关系。本文完全在模拟中隔离了产生该最优的因素,并描绘了其成立的范围。(1) 现象:保持能力的倒U型出现在多个相关任务持续学习基准上(Split-MNIST、FashionMNIST、持续阴阳)。 (2) 隔离:幅度匹配的阶梯显示,该效应需要向整合权重进行相干恢复——相同幅度的随机方向力不产生最优,而指向错误目标的相干力则主动造成损害。 (3) 活性成分:通过将锚定增益与注入噪声方差sigma^2耦合,可恢复大部分最优——这是一行规则,而Ornstein-Uhlenbeck自适应(固定增益)和MESU(后验方差增益)均未实现。强制Ornstein-Uhlenbeck计算推导出上升沿,并预测最优噪声随每任务干扰g增加而上升——在样本外方向上与先前测量结果一致(指数在我们的网格上未解决)。源自Doob h变换的势垒调节是低sigma安全网,在耦合增益过弱时限制遗忘。 (4) 适用范围:该最优需要共享任务结构——在permuted-MNIST上不存在,受控的旋转与置换比较将边界定位到任务结构;精确的控制量留待未来研究。 (5) 长度:在匹配严重程度下,优势持续但随任务数量减弱,我们证明没有旋转族能归因该趋势(紧致群恒等式)。原始规则在BrainScaleS-2上的单种子演示另行报告(Howe, arXiv:2607.06924);本文不涉及硬件声明。

英文摘要

Injecting stochastic noise into a consolidation rule can improve a network's retention of earlier tasks up to an optimal level, then degrade it -- an inverted-U in retention vs. noise. This paper isolates what produces that optimum and maps where it holds, entirely in simulation. (1) Phenomenon: the retention inverted-U appears on several related-task continual-learning benchmarks (Split-MNIST, FashionMNIST, continual Yin-Yang). (2) Isolation: a magnitude-matched ladder shows the effect requires coherent restoring toward the consolidated weights -- a random-direction force of identical magnitude produces no optimum, and a coherent force toward the wrong target actively hurts. (3) Active ingredient: most of the optimum is recovered by coupling the anchor gain to the injected-noise variance sigma^2 -- a one-line rule that neither Ornstein-Uhlenbeck Adaptation (fixed gain) nor MESU (posterior-variance gain) implements. A forced Ornstein-Uhlenbeck calculation derives the rising flank and predicts that the optimal noise rises with per-task interference g -- confirmed out-of-sample in direction against pre-existing measurements (the exponent is unresolved at our grid). The barrier-conditioning of the originating Doob h-transform is a low-sigma safety net that bounds forgetting where the coupled gain is too weak. (4) Scope: the optimum requires shared task structure -- it is absent on permuted-MNIST, and a controlled rotated-vs-permuted comparison localizes the boundary to task structure; the precise governing quantity is left open. (5) Length: at matched severity the advantage persists but attenuates with task count, and we show no rotation family can attribute the trend (a compact-group identity). A single-seed BrainScaleS-2 demonstration of the originating rule is reported separately (Howe, arXiv:2607.06924); this paper makes no hardware claim.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑