arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33436cs.LG

SchemaMem:用于延迟状态检索的架构索引循环记忆

SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval

Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung

首次发表
浏览论文内容

中文总结 AI 辅助

SchemaMem是一种结合块局部注意力和架构索引相位状态的循环记忆架构,在延迟状态检索任务中实现了更高的写入值保留率,但需要更多优化步骤,揭示了保留与优化之间的权衡。

中文摘要 AI 辅助

注意力机制提供了对过去表征的直接访问,但保留不断增长的历史记录代价高昂。循环模型限制了持久状态的大小,却必须在处理后续输入时保留选定的信息。我们提出了SchemaMem,一种基于注意力的循环记忆架构,结合了块局部注意力与持久的、架构索引的相位状态。学习得到的架构嵌入为读写操作提供了共享的表征参考。读取使用当前状态,而写入使用层输入和静态架构嵌入,排除了来自该层自身状态的直接反馈。块边界提交通过前向计算聚合有界的相位增量。相同的参数也支持在循环训练之前和期间进行全历史注意力训练。我们在一个受控的地址-值任务中研究了选择性更新、保留和延迟检索,比较了具有大致匹配的参数数量和持久状态维度的三层模型。在九个地址/值设置和三个训练种子下,SchemaMem在最大训练延迟四倍处的平均写入值保留率高于两个基线,而这两个基线则被训练以达到更高的范围内精度目标。更新值恢复在全部九个设置中有利于SchemaMem对抗Mamba-3,在七个设置中对抗Gated DeltaNet。默认设置在延迟处一致地有利于Gated DeltaNet而非SchemaMem,且SchemaMem需要显著更多的优化步骤。这些结果揭示了架构索引循环中一个有前景的保留-优化权衡。

英文摘要

Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-based recurrent memory architecture combining chunk-local attention with a persistent, schema-indexed phase state. Learned schema embeddings provide a shared representational reference for reading and writing. Reads use the current state, whereas writes use the layer input and static schema embeddings, excluding direct feedback from that layer's own state. Chunk-boundary commits aggregate bounded phase increments through forward computation. The same parameters also support full-history attention training before and during recurrent training. We studied selective updates, preservation, and delayed retrieval in a controlled address--value task, comparing three-layer models with approximately matched parameter counts and persistent-state dimensions. Across nine address/value settings and three training seeds, SchemaMem has higher mean written-value retention at four times the maximum training delay than both baselines, which are trained toward a higher in-range accuracy target. Updated-value recovery favors SchemaMem in all nine settings against Mamba-3 and seven against Gated DeltaNet. Defaults consistently favor Gated DeltaNet over SchemaMem at that delay, and SchemaMem requires substantially more optimization steps. These results identify a promising retention--optimization trade-off in schema-indexed recurrence.

发表机构

  • Chungnam National University(忠南国立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑