arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

固定状态递归中联想记忆的解剖:匹配状态分解、干扰墙与打破它的课程

Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

Julian Boesch, Andrew Wee

arXiv 2609.16183首次发表:更新:

发表机构

Purdue University; Obit Research(普渡大学; Obit Research)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过匹配状态分解,发现卷积主导联想记忆,秩1转移优势在匹配后消失;距离课程和精度门控斜坡可突破干扰墙,将锁定率从1/10提升至7/10。

AI 中文摘要

固定状态递归——线性注意力和状态空间模型——据报道在联想记忆方面落后于注意力机制,但整体架构的比较无法说明是哪个成分导致了这种差距。我们在固定状态预算下,沿三个单旋钮轴分解掩蔽多查询记忆:一个短因果卷积、转移结构(秩1 delta规则 vs. 对角)以及衰减。卷积占主导地位(在匹配训练下,两个家族中约+0.5的记忆提升):将无卷积单元与带卷积的Mamba进行比较,测量的是缺失的卷积,而非递归本身。秩1转移在16/32对时比其对消融高+0.19/+0.32,但一旦两个单元都带有卷积,差距缩小到+0.03,且状态匹配的Mamba-2与无武装的秩1单元持平:没有类别主张能幸存。能解决32对记忆的单元在负载增加时优雅地退化,但在从干扰物草堆中检索4对时却降至随机水平——在长度和转移上均平坦。干扰是在稀疏监督下产生的,而非容量问题:一个距离课程将未改变的架构从0.021提升到1.000。训练是一场锁定彩票——一个种子要么锁定要么不锁定——而课程是杠杆。锁定率从1/10上升到7/10(p=0.02);密集监督没有增加任何东西;在L=256时,一个成形斜坡重新打开了统一课程无法打开的边界(4/5 vs. 0/9);在L=512时,斜坡崩溃(0/6),而根据测量精度对其进行门控则锁定6/6(p=0.001)。双向去噪单元在读取查询之前先读取草堆,与因果训练相比没有显示出可测量的优势(十个种子),且碰撞键检索需要两层。在S_5状态跟踪护栏上,武装记忆是免费的——武装单元在每个深度都显著更好(p<=0.0044)。这些结果用测量分解和两个廉价干预取代了“递归模型在记忆方面表现不佳”的说法。

英文摘要

Fixed-state recurrences--linear attention and state-space models--are reported to lag behind attention on associative recall, but whole-architecture comparisons cannot say which ingredient is responsible. We decompose masked multi-query recall at a fixed state budget along three single-knob axes: a short causal convolution, the transition structure (rank-1 delta rule vs. diagonal), and decay. The convolution dominates (~+0.5 recall in both families under matched training): comparisons that pit convolution-free cells against a convolution-equipped Mamba measure the missing convolution, not the recurrence. The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs, but the margin shrinks to +0.03 once both cells carry the convolution, and a state-matched Mamba-2 ties the unarmed rank-1 cell: no class claim survives. Cells that solve 32-pair recall degrade gracefully with load yet fall to chance retrieving 4 pairs from a distractor haystack--flat across lengths and transitions. Interference under sparse supervision, not capacity: a distance curriculum takes the unchanged architecture from 0.021 to 1.000. Training is a lock-in lottery--a seed either locks in or does not--and the curriculum is the lever. Lock-in rises from 1/10 to 7/10 (p=0.02); dense supervision adds nothing; at L=256 a shaped ramp reopens a boundary the uniform curriculum cannot (4/5 vs. 0/9); and at L=512, where the ramp collapses (0/6), gating it on measured accuracy locks in 6/6 (p=0.001). Bidirectional denoiser cells, reading the query before the haystack, show no measurable advantage over causal training (ten seeds), and collision-key retrieval needs two layers. Arming for recall is free on an S_5 state-tracking guardrail--the armed cell is significantly better at every depth (p<=0.0044). These replace "recurrent models are bad at recall" with a measured decomposition and two cheap interventions.

Comments14 pages, 2 figures, 6 tables. Preprint of preliminary results; code and result JSONs at https://github.com/JIBSIL/dualgoose

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑