arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

最新精确匹配注意力

Latest Exact Match Attention

Moritz Brösamle

arXiv 2609.25802首次发表:更新:

发表机构

University of Tübingen(蒂宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出最新精确匹配注意力(LEMA),通过二值化查询键和精确匹配实现高效Transformer,理论证明与word-RAM等价,训练方法有效,在联想回忆和语言建模上表现良好。

AI 中文摘要

我们引入了最新精确匹配注意力(LEMA),这是一种Transformer的注意力变体,其中查询和键被二值化,每个查询仅关注最新的精确匹配键。我们证明了具有思维链的LEMA Transformer可以模拟word-RAM,正如最近针对限制较少的右端硬注意力所展示的那样。与先前的硬注意力变体相比,对精确匹配的限制使得高效的反向方向成为可能:word-RAM可以模拟LEMA Transformer,每个token的成本与上下文长度无关。这些结果共同表明,两种计算模型在计算和内存方面存在紧密对应关系。在理论之外,我们提出了一种针对LEMA Transformer的训练方法,该方法通过直通估计器处理其二值化的不可微操作,并通过软注意力替代物退火至LEMA。在合成联想回忆任务中,以这种方式训练的LEMA模型利用其增长的状态来存储和回忆大量关联,优于具有固定状态大小的门控DeltaNet(GDN)。作为首次扩展测试,我们训练了参数多达8.34亿的LEMA语言模型。它们在损失上匹配约为其一半大小的softmax Transformer,并且在重复稀有短语和针检索任务上,虽然落后于softmax Transformer,但在更长距离上的回忆能力优于同等规模的GDN模型。最后,我们实现了基于字典的LEMA Transformer推理,并展示了与GDN相当的恒定生成速度,尽管其状态不断增长,字典驻留在主内存而非VRAM中。代码可在该https URL获取。

英文摘要

We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends only to the latest exactly matching key. We prove that LEMA transformers with chain of thought can simulate word-RAMs, as was recently shown for the less restrictive rightmost hard attention. In contrast to prior hard attention variants, the restriction to exact matches enables an efficient converse direction: word-RAMs can simulate LEMA transformers at a cost per token independent of the context length. Together, these results yield a close correspondence between the two computational models in terms of both compute and memory. Beyond the theory, we propose a training method for LEMA transformers that handles their non-differentiable operations with a straight-through estimator for the binarization and a soft attention surrogate annealed towards LEMA. On a synthetic associative recall task, LEMA models trained this way use their growing state to store and recall a large number of associations, outperforming gated DeltaNet (GDN) with its fixed state size. As a first scaling test, we train LEMA language models with up to 834 million parameters. They match softmax transformers of around half their size in loss and, on repeated rare phrases and a needle-retrieval task, remain behind softmax transformers but recall across longer distances than GDN models of comparable size. Finally, we implement dictionary-based inference for LEMA transformers and show constant generation speed comparable to GDN despite their growing state, with the dictionaries residing in main memory rather than VRAM. Code is available at https://github.com/moritzbroe/latest_exact_match_attention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑