arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TransMem:将隐藏状态转换为大语言模型的记忆

TransMem: Transforming Hidden States into Memory for Large Language Models

Haodong Lei, Junming Liu, Yirong Chen, Pinlong Cai, Botian Shi, Ding Wang, Hongsong Wang

arXiv 2607.29032首次发表:更新:

发表机构

Southeast University; Shanghai Artificial Intelligence Laboratory(东南大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长上下文LLM智能体的历史信息未充分利用问题,提出轻量级记忆模块TransMem,结合证据条件自蒸馏,在多基准测试中显著提升性能。

AI 中文摘要

大语言模型(LLM)智能体越来越多地在长交互历史上运行,有效推理需要识别和利用分布在过去观测与动作中的任务相关证据。然而,先前计算的表示中编码的有用信息在后续生成过程中常未被充分利用。我们提出**TransMem**,这是一个轻量级推理时参数化记忆模块,它将来自冻结LLM主干的稀疏历史隐藏状态转换为可重用的记忆表示。TransMem使用轻量级门控网络,动态地将潜在干预应用于当前隐藏状态,无需重复编码先前上下文。为学习可迁移的记忆利用而非任务特定知识,我们引入证据条件自蒸馏:记忆增强的学生模型处理完整上下文,并匹配仅使用证据的教师模型的预测分布,二者共享同一冻结主干。在LoCoMo、HotpotQA和MemoryAgentBench上的实验表明,在不同模型架构和规模上均有一致提升:TransMem在LoCoMo上获得11.58--29.25的$F_1$值提升,在HotpotQA上获得10.20--13.03的$F_1$值提升,同时将MemoryAgentBench的平均准确率从29.54%提高到40.00%。这些结果确立了稀疏历史隐藏状态是长上下文LLM智能体的有效且高效的记忆底物,代码可在该https URL获取。

英文摘要

Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distributed across past observations and actions. However, useful information encoded in previously computed representations is often underutilized during subsequent generation. We propose \textbf{TransMem}, a lightweight inference-time parametric memory module that transforms sparse historical hidden states from a frozen LLM backbone into reusable memory representations. TransMem uses a lightweight gating network to dynamically apply the latent intervention to the current hidden states, without repeatedly encoding the preceding context. To learn transferable memory utilization rather than task-specific knowledge, we introduce evidence-conditioned self-distillation. A memory-augmented student processes the full context and matches the predictive distribution of an evidence-only teacher that shares the same frozen backbone. Experiments on LoCoMo, HotpotQA, and MemoryAgentBench demonstrate consistent improvements across different model architectures and scales. TransMem yields gains of 11.58--29.25 $F_1$ on LoCoMo and 10.20--13.03 $F_1$ on HotpotQA, while improving the average MemoryAgentBench accuracy from 29.54\% to 40.00\%. These results establish sparse historical hidden states as an effective and efficient memory substrate for long-context LLM agents. Our code is available at https://github.com/Haodong-Lei-Ray/TransMem.

Comments12 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑