arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在持续学习中通过可学习小波激活维持可塑性

Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning

Zeyang Zhang, Tieliang Gong, Junyan Lu, Weizhan Zhang

arXiv 2608.12874首次发表:更新:

发表机构

Institute of Multimedia Knowledge Fusion and Engineering(多媒体知识融合与工程研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对持续学习的可塑性损失与谱偏差问题,提出可学习小波激活及动态小波注入等方法,在多基准上实现了先进性能。

AI 中文摘要

可塑性损失是持续学习中的关键挑战,严重阻碍了序列任务的获取。优化激活设计虽为潜在解决方案,但当前固定形式函数存在固有的低频谱偏差,而可学习变体则允许无约束更新,易引发灾难性遗忘。为解决这些局限,我们提出一种新型可学习小波激活,将激活函数分解为低频与高频分量以明确对抗谱偏差;同时采用动态小波注入,自适应增强新任务的可塑性,并辅以正则化策略确保先前所学知识的稳定性。理论上,我们为该框架提供了严格数学保证,证明混合小波架构对高效$L^2$近似的结构必要性,且解耦学习率机制可成功恢复网络对高频信息的可塑性;此外还对损失驱动注入触发机制进行形式推导,以精准引导注入。大量实证评估表明,我们的方法在整个学习过程中保持优异的可训练性与泛化能力,在各类持续学习基准上达到了先进性能。

英文摘要

Plasticity loss has emerged as a critical challenge in continual learning that significantly hinders the acquisition of sequential tasks. While optimizing activation designs offers a potential solution, current fixed-form functions suffer from an inherent spectral bias towards low-frequency variations, whereas learnable variants permit unconstrained updates that induce catastrophic forgetting. To address these limitations, we propose a novel learnable wavelet activation that decomposes the activation function into low-frequency and high-frequency components to explicitly counter spectral bias. Furthermore, we employ dynamic wavelet injection to adaptively enhance plasticity for new tasks, alongside a regularization strategy to ensure the stability of previous learned knowledge. Theoretically, we provide rigorous mathematical guarantees for the proposed framework, proving the structural necessity of the hybrid wavelet architecture for efficient $L^2$ approximation and demonstrating that the decoupled learning rate mechanism successfully restores network plasticity for high-frequency information. Additionally, we provide a formal derivation of the loss-driven injection trigger mechanism to precisely guide the injection. Extensive empirical evaluations demonstrate that our approach maintains superior trainability and generalization throughout the learning process and achieves state-of-the-art performance across diverse continual learning benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑