arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33268cs.AI

LSTMem:面向大型语言模型的分层长短期在线记忆

LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

Xianglong Shi, Ruijie Yang, Sirui Zhao, Shukang Yin, Zihao Bian, Tinghao Yi, Enhong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM长期记忆问题,提出LSTMem,通过分离细胞状态与隐藏状态并分层传播,在多个基准上提升记忆性能。

中文摘要 AI 辅助

大型语言模型日益充当长期助手和智能体,它们既需要在交互中积累信息,又需要在后续请求依赖这些信息时提供相关部分。现有的紧凑型在线记忆通常使用单一持久状态来同时积累历史和提供读出,因此记忆存储的内容无法与当前计算所暴露的内容分开控制。我们提出LSTMem,一种受LSTM启发的在线记忆,它为冻结的LLM的每一层配备两个矩阵值状态:一个积累历史的细胞状态和一个其读出校正主干注意力的隐藏状态。输入门和遗忘门控制细胞存储的内容,而输出门则单独控制细胞通过隐藏状态暴露的内容。LSTMem还通过前向隐藏状态传播和块末反馈在深度上连接记忆,并使用高层重建梯度来细化低层细胞状态,然后从浅层到深层重建隐藏状态。在Qwen3-4B-Instruct上的记忆基准测试中,LSTMem在MemoryAgentBench、LoCoMo和HotpotQA上持续优于普通主干。比较进一步表明,基于LSTM的记忆公式优于关联记忆对应物,而移除跨层隐藏记忆传播会降低性能。这些结果证明了将记忆积累与记忆表达分离以及跨模型深度分层组织记忆的优势。代码可在以下网址获取:此https URL。

英文摘要

Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to accumulate history and to serve readout, so what the memory stores cannot be controlled separately from what it exposes to the current computation. We propose LSTMem, an LSTM-inspired online memory that instead equips each layer of a frozen LLM with two matrix-valued states: a cell state that accumulates history and a hidden state whose readouts correct the backbone's attention. Input and forget gates control what the cell stores, while an output gate separately controls what the cell exposes through the hidden state. LSTMem further connects memory across depth through forward hidden-state propagation and block-end feedback, and uses higher-layer reconstruction gradients to refine lower-layer cell states before rebuilding hidden states from shallow to deep layers. Across memory benchmarks on Qwen3-4B-Instruct, LSTMem consistently improves MemoryAgentBench, LoCoMo, and HotpotQA over the plain backbone. Comparisons further show that the LSTM-based memory formulation outperforms an associative-memory counterpart, while removing cross-layer hidden-memory propagation degrades performance. These results demonstrate the benefits of separating memory accumulation from memory expression and organizing memory hierarchically across model depth. The code is available at https://github.com/Longchentong/LSTMem.

↑