发表机构
Frederick University(弗雷德里克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对无注意力序列模型的长程单次回忆缺陷,提出带特定机制的“笔记本”记忆单元,在合成与真实文本任务上实现长度不变的精确回忆,且可复现。
AI 中文摘要
循环无注意力序列模型存在结构缺陷:衰减的状态无法精确回忆过去久远的单次所见内容。我们为Kathleen主干模型添加了第二层记忆单元——“笔记本”:一个固定键的全息(HRR)关联存储,带有学习到的局部写入门控、自门控原始读取和写入触发遗忘,共25K参数,可附加到任意主干的logits上。(1)机制:在受控的“干草堆中找针”任务上,笔记本在训练长度4倍时达到80-82%的一次性回忆准确率,而裸主干得分约4%,参数匹配的注意力头在训练长度内得100%、超出则得0%;该存储的寻址天生具有长度不变性,未训练的记忆单元在512、2048和4096字节下的准确率均达90%;由于存储是线性叠加,仅通过算术就衍生出两种能力:选择性遗忘(一次减法可将一个事实擦除至随机水平,保留的事实不受影响)和逐词归因(反事实擦除可命名每个正确字节的源事实,实现100%溯源)。(2)真实文本:在WikiText-2字节级任务上,笔记本提升了重复稀有词的预测性能,每字节增益为+0.15-0.27比特,增益随提及间隔增大而上升,在训练长度4倍时保持零样本性能;写入触发遗忘在8倍长度时消除了内存污染(首次提及的代价从+0.33降至-0.004)。(3)范围与规模:参数匹配的注意力头确实能在自然文本重复上泛化,因此笔记本的核心是在O(L)复杂度下实现精确回忆;在WikiText-103阶梯数据集(8至512MB)上,零样本重复增益单调上升;所有实验均已预注册,报告了随机种子,可在单个免费层级GPU上复现。
英文摘要
Recurrent, attention-free sequence models share a structural weakness: a fading state cannot perform exact recall of something seen once, far in the past. We add to the Kathleen trunk a second memory layer -- a "notebook": a fixed-key holographic (HRR) associative store with a learned local write gate, a self-gating raw read, and write-triggered forgetting -- 25K parameters that attach to the logits of any trunk. (1) Mechanism: on a controlled needle-in-haystack task the notebook reaches 80-82% one-shot recall at 4x the training length, where the bare trunk scores ~4% and a parameter-matched attention head scores 100% inside its training length and 0% beyond it. Addressing is length-invariant by construction; the untrained memory alone recalls at 90% accuracy identically at 512, 2048 and 4096 bytes. Because the store is a linear superposition, two capabilities follow from arithmetic alone: selective unlearning (one subtraction erases one fact to chance, retained facts unharmed) and per-token attribution (counterfactual erasure names the source fact of every correct byte, 100% provenance). (2) Real text: on WikiText-2 bytes the notebook improves prediction of repeated rare words by +0.15-0.27 bits/byte, the gain growing with the distance between mentions and holding zero-shot at 4x training length; write-triggered forgetting eliminates memory pollution at 8x length (first-mention cost +0.33 -> -0.004). (3) Scope and scale: a parameter-matched attention head does generalize on natural-text repetition, so the notebook's claim is exact recall at O(L); on a WikiText-103 ladder (8 to 512 MB) the zero-shot repeat gain rises monotonically. All experiments are pre-registered, seeds reported, and reproducible on a single free-tier GPU.
Comments12 pages, 3 figures. Paper 4 of the Kathleen series. Mechanism covered by U.S. Provisional Patent Application No. 64/140,260