arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19521cs.LG

LSTM-UT 与循环深度 Transformer 在元胞自动机上的研究

Storing Is Not Remembering: LSTM-UT and Bounded Gated Memory for Looped Transformers

Aras Kavuncu, Muhammad Burhan Hafez

首次发表
浏览论文内容

中文总结 AI 辅助

本研究比较了三种循环深度 Transformer 在元胞自动机上的表现,提出 LSTM-UT 模型,其有界门控记忆有效提升了深度外推和延迟召回能力。

中文摘要 AI 辅助

循环深度 Transformer 重复应用共享计算,但在跨步骤保留信息的方式上有所不同。我们比较了块通用 Transformer(BUT),它仅携带当前隐藏状态;CoTFormer,它还保留一个不断扩展的注意力缓存;以及一种具有有界门控记忆的新型 LSTM 通用 Transformer(LSTM-UT)。在 Rule 30 元胞自动机上,BUT 比 CoTFormer 更可靠地外推到未见过的循环深度,尽管其准确性最终会下降。状态和缓存干预表明,CoTFormer 的失败取决于它们的交互:纠正当前状态可以暂时恢复准确性,而保留的历史信息可能破坏这种纠正。在延迟召回任务中,尽管缺乏对过去状态的直接访问,BUT 也优于 CoTFormer;CoTFormer 不能可靠地选择所请求的缓存表示。LSTM-UT 在这些基线上改善了深度外推和延迟召回。结果支持有界门控记忆作为这些任务中重复计算和后续检索的有效归纳偏置。

英文摘要

Recurrent-depth Transformers reuse one block across many steps, so information needed later must survive repeated rewriting of the hidden state. A natural remedy is to keep more history. We show that, in controlled cellular-automaton tasks, making history available is not the same as making it usable. Using Rule 30, where the correct state is known at every recurrent step, we test depth extrapolation and de- layed recall, the recovery of an earlier state after further computation. CoTFormer, which caches keys and values from every earlier step, extrapolates less far and recalls less accurately than a Block Universal Transformer (BUT) that keeps only its current state. Interventions show that its retained history can pull a corrected trajectory back toward failure, and that the cache block written at the requested step is neither necessary nor sufficient for recall. We introduce LSTM-UT, which adds a small, bounded, gated cell state to the shared block. Trained to depth 12, LSTM-UT keeps 99.7% exact-row accuracy at depth 60, where BUT gets no row fully correct, and one checkpoint stays above 99.95% at depth 1,000. It also improves delayed recall over both baselines, and the advantage largely persists at near-matched parameter counts. On these tasks, a small state under learned control proved more useful than a complete but unaddressed history. In OpenWebText2 language modelling, LSTM-UT outperforms BUT and, at equal width, reaches slightly lower perplexity than CoTFormer while CoTFormer needs up to 91% more training time per step; against a parameter-matched CoTFormer, LSTM-UT comes within 0.6 perplexity.

发表机构

  • University of Southampton(南安普顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑