arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28399cs.LG

记忆注意力

Memory Attention

Jiale Kang

首次发表
浏览论文内容

中文总结 AI 辅助

提出记忆注意力(MA)方法,用令牌索引记忆替代值投影并结合上下文键构建注意力值,在匹配训练预算下提升语言建模与下游性能。

中文摘要 AI 辅助

语言模型通常从上下文隐藏状态构建注意力值,即使其中部分内容可能在不同上下文中可重用。我们研究了当辅以上下文信息时,以令牌索引的记忆能否替代专用的值投影。我们提出了记忆注意力(Memory Attention, MA),该方法通过结合层特定的令牌记忆与上下文键来形成值。记忆提供令牌特定的表示,而键保留上下文依赖性。在推理时,归一化可以折叠进记忆表中,将值的构建简化为查找和加法操作。令牌索引检索还支持带预取的CPU卸载,从而减少GPU参数存储。在匹配的训练令牌预算下并增加额外记忆参数时,跨注意力配置的实验表明语言建模和平均下游性能均有提升。

英文摘要

Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and addition. Token-indexed retrieval also enables CPU offloading with prefetching, reducing GPU parameter storage. Under matched training token budgets and with additional memory parameters, experiments across attention configurations show improved language modeling and average downstream performance.

↑