发表机构
The University of Alabama(阿拉巴马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对视觉Transformer难以在不同时间尺度组织新信息的问题,提出分层赫布记忆架构,在Omniglot、CORe50任务上取得优于基线的识别与关联准确率,验证了其组织在线视觉经验的有效性。
AI 中文摘要
视觉Transformer能提供强大的视觉表征,但通常依赖缓慢更新的参数,限制了其在不同记忆时间尺度上组织新获取信息的能力。本研究提出了分层赫布记忆(Hierarchical Hebbian Memory),这是一种由快速工作记忆(Working Memory)、持久路由情景记忆(Routed Episodic Memory)和较慢的语义记忆(Semantic Memory)组成的三级记忆架构。一个学习得到的控制器调节记忆贡献、读写路由、可塑性、保留和巩固过程。因果读写前生命周期确保当前结果不会影响它所监督的预测。该架构在Omniglot 5-way 1-shot识别和CORe50持续物体识别任务上进行了评估。结合经验回放(experience replay)时,使用Swin-Tiny的分层模型在Omniglot上达到97.39%的准确率,在CORe50上达到95.37%的最终准确率。学习得到的多库检索达到47.50%的延迟关联准确率,而单个持久库为24.17%,无记忆时为25.00%。在干扰项介入后,情景记忆与存储关联的余弦相似度约为0.96,而工作记忆降至约0.05。这些结果表明,赫布关联与学习得到的记忆路由可共同在视觉Transformer内部的快速、持久和巩固的记忆时间尺度上组织在线视觉经验。
英文摘要
Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, limiting their ability to organize newly acquired information across different memory timescales. This work proposes \textit{Hierarchical Hebbian Memory}, a three-level memory architecture composed of rapid Working Memory, persistent Routed Episodic Memory, and slower Semantic Memory. A learned controller regulates memory contribution, read and write routing, plasticity, retention, and consolidation. A causal read-before-write lifecycle ensures that the current outcome cannot influence the prediction it supervises. The architecture is evaluated on Omniglot 5-way 1-shot recognition and CORe50 continual object recognition. With Swin-Tiny, the hierarchical model reaches 97.39\% accuracy on Omniglot and 95.37\% final accuracy on CORe50 when combined with experience replay. Learned multi-bank retrieval reaches 47.50\% delayed-association accuracy, compared with 24.17\% for a single persistent bank and 25.00\% without memory. After intervening distractors, Episodic Memory retains approximately 0.96 cosine similarity with stored associations, while Working Memory falls to approximately 0.05. These results show that Hebbian association and learned memory routing can jointly organize online visual experience across rapid, persistent, and consolidated memory timescales within Vision Transformers.