arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回忆录:模型在思考时应该写入其内存吗?

Memoir: Should a Model Write to Its Memory While It Thinks?

Jaber Jaber, Osama Jaber

arXiv 2607.20792首次发表:更新:

发表机构

RightNow AI(即时人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨模型思考时写入内存的情况,通过结合多种技术构建回忆录模型,对比耦合与只读思考臂在程序关联回忆任务中的表现,发现耦合臂在训练初期有学习速度惩罚,未出现内存重写破坏能量信号的情况,内核重组减少了前向时间。

AI 中文摘要

回忆录结合了每个样本的快速内存、共享的慢速参数、可变深度的潜在递归以及未来潜在能量目标。我们测试了其最具风险的耦合:每次思考迭代可能会重写同一迭代读取的快速层。在具有关键干扰的程序关联回忆中,我们将一个耦合臂与一个相同的只读思考臂进行比较。两个臂都包含81,738个参数,包括76,362个可训练参数,并使用匹配的声明前向乘积累加计数、数据、优化器、调度和种子。经过12个种子的240个训练步骤后,耦合回忆为0.5203,95%区间为[0.4522, 0.5883],而只读回忆为0.6557,区间为[0.5953, 0.7160]。臂按种子配对,只读领先0.1354,在11个自由度上的配对t为3.23,差异的95%区间为[0.0431, 0.2277],在12个种子中的10个上获胜。经过8个种子的960个步骤后,两个臂都达到了1.0000,所以测量到的效果是在固定预算下的学习速度惩罚,而不是已证明的能力惩罚。更长的控制受上限限制,未测量在非饱和任务上的收敛情况。预测的内存重写会破坏能量信号的失败并未发生:能量余量增加并保持。内核重组还将指定设备上的增量规则前向时间从0.907毫秒减少到0.351毫秒。代码和证据可在此https URL获取。

英文摘要

Memoir combines per-sample fast memory, shared slow parameters, variable-depth latent recurrence, and a future-latent energy objective. We test its riskiest coupling: each pondering iteration may rewrite the fast tier that the same iteration reads. On procedural associative recall with key interference, we compare a coupled arm against an otherwise identical read-only pondering arm. Both arms contain 81,738 parameters, including 76,362 trainable parameters, and use matched declared forward multiply-accumulate counts, data, optimizer, schedule, and seeds. After 240 training steps across 12 seeds, coupled recall is 0.5203 with a 95 percent interval of [0.4522, 0.5883], while read-only recall is 0.6557 with [0.5953, 0.7160]. The arms are paired per seed, and the read-only lead of 0.1354 gives a paired t of 3.23 on 11 degrees of freedom with a 95 percent interval of [0.0431, 0.2277] on the difference, winning on 10 of 12 seeds. After 960 steps across 8 seeds, both arms reach 1.0000, so the measured effect is a learning-speed penalty at a fixed budget, not a demonstrated capability penalty. That longer control is ceiling limited, leaving convergence on a non-saturating task unmeasured. A predicted failure in which memory rewriting corrupts the energy signal did not occur: the energy margin grew and held. Kernel restructuring also reduced delta-rule forward time from 0.907 ms to 0.351 ms on the stated device. Code and evidence are available at https://github.com/RightNow-AI/Memoir

Comments9 pages, 4 figures, 4 tables, 1 algorithm. Code: https://github.com/RightNow-AI/Memoir

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑