arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28663cs.NEcs.LG

NeuroSynth:一种受生物启发的持续强化学习架构,用于缓解灾难性遗忘

NeuroSynth: A Biologically Inspired Continual Reinforcement Learning Architecture for Mitigating Catastrophic Forgetting

Yash Kini

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出受生物启发的持续强化学习架构NeuroSynth,通过双路径巩固机制缓解灾难性遗忘,实验显示其在多个顺序导航任务中相比PPO、EWC保留更多早期任务知识,改善了持续学习的稳定性-可塑性平衡。

中文摘要 AI 辅助

人工智能(AI)系统通常在孤立任务上表现良好,但在持续学习场景中表现不佳,在该场景下对新任务进行训练会覆盖先前获取的知识,这种失效模式被称为灾难性遗忘。生物学习系统通过互补记忆过程减少这种干扰,该过程涉及快速的海马编码和较慢的皮层巩固。本研究引入NeuroSynth,一种受大脑启发的持续强化学习架构,旨在通过双路径巩固机制缓解灾难性遗忘。NeuroSynth使用不同的“计划”和“习惯”路径,结合重放和知识蒸馏,将快速任务获取与长期保留分离开来。在非重访持续学习场景中,针对三个目标位置变化的顺序导航任务,将NeuroSynth与近端策略优化(PPO)和弹性权重巩固(EWC)进行了评估。在六个独立随机种子下,顺序训练后,NeuroSynth保留的早期任务知识远多于PPO,任务A成功率为18.00%,而PPO为0.33%(p=0.014929,Cohen's d=1.49);任务B成功率为35.33%,而PPO为0.00%(p=0.002376,Cohen's d=2.31)。NeuroSynth的最终任务C性能也高于EWC,达到9.00%,而EWC为2.00%(p=0.226643,Cohen's d=0.56),表明存在中等但无统计学意义的优势。这些发现表明,受生物启发的巩固机制可能会改善持续强化学习系统中的稳定性-可塑性平衡。

英文摘要

Artificial Intelligence (AI) systems often perform well on isolated tasks but struggle under continual learning conditions, where training on new tasks can overwrite previously acquired knowledge, a failure mode known as catastrophic forgetting. Biological learning systems reduce this interference through complementary memory processes involving rapid hippocampal encoding and slower cortical consolidation. This study introduces NeuroSynth, a brain-inspired continual reinforcement learning architecture designed to mitigate catastrophic forgetting through a dual-pathway consolidation mechanism. NeuroSynth separates rapid task acquisition from long-term retention using distinct "plan" and "habit" pathways combined with replay and knowledge distillation. NeuroSynth was evaluated against Proximal Policy Optimization (PPO) and Elastic Weight Consolidation (EWC) across three sequential navigation tasks with changing goal locations in a non-revisitation continual learning setting. Across six independent seeds, NeuroSynth preserved substantially more early-task knowledge than PPO after sequential training, achieving 18.00% Task A success rate compared to 0.33% for PPO (p = 0.014929, Cohen's d = 1.49) and 35.33% Task B success rate compared to 0.00% for PPO (p = 0.002376, Cohen's d = 2.31). NeuroSynth also demonstrated higher final Task C performance than EWC, achieving 9.00% compared to 2.00% (p = 0.226643, Cohen's d = 0.56), indicating a moderate but not statistically significant advantage. These findings suggest that biologically inspired consolidation mechanisms may improve the stability-plasticity balance in continual reinforcement learning systems.

发表机构

  • James Madison High School(詹姆斯麦迪逊高中)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑