arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RPMem:面向LLM智能体的跨会话长期循环参数化记忆学习

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

Fanyu Zhao, Ruike Cao, Liang Dong, Fugen Yao, Jian Xu, Guanjun Jiang, Han Zhang, Yifei Zhao, Yinsheng Li

arXiv 2609.23466首次发表:更新:

发表机构

Fudan University; Qwen Applications Business Group, Alibaba Group(复旦大学; 阿里巴巴集团通义千问应用业务集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RPMem提出两阶段架构,将会话编译为模型无关的潜在记忆并经循环门控跨会话整合,映射为LoRA参数,实现跨骨干可复用的演化记忆,在多个基准上超越现有方法。

AI 中文摘要

长时间运行的LLM智能体需要跨会话持久存在并不断演化的记忆。基于文本的记忆在每次查询时检索并重建过去的交互,使得随着历史增长,长期性能越来越依赖于检索质量和上下文推理。参数化记忆将经验直接编码到模型计算中,但现有方法对跨会话记忆演化的支持有限。它们与特定骨干模型的耦合进一步限制了模型替换后记忆的复用。我们提出了RPMem,一种两阶段架构,通过前向计算将每个会话编译为与模型无关的潜在记忆,并通过任务训练的循环门控将其与保留记忆选择性整合。整合后的记忆随后映射到骨干特定的低秩适配(LoRA)参数,使得编码能力在骨干替换时得以迁移。在三个长期记忆基准和五个不同骨干上的评估展示了广泛的泛化能力,且更新成本和内存占用近乎恒定。在PERMA上使用Qwen3-8B,RPMem达到85.52%,分别超过最强的参数化和基于文本的基线5.32和12.98个百分点。消融实验验证了会话编译和跨会话整合的互补作用,而动态分析揭示门控获得了任务特定的记忆整合策略。这些结果确立了RPMem作为一个生命周期无关的参数化记忆框架,能够维护跨会话演化的记忆,并在骨干替换后保持可复用性。我们的实现可在该https URL获取。

英文摘要

Long-running LLM agents require memory that persists and evolves across sessions. Text-based memory retrieves and reconstructs past interactions at every query, making long-horizon performance increasingly dependent on retrieval quality and contextual reasoning as histories grow. Parametric memory encodes experience directly into model computation, but existing approaches provide limited support for cross-session memory evolution. Their coupling to a specific backbone further restricts memory reuse after model replacement. We introduce RPMem, a two-stage architecture that compiles each session into a model-independent latent memory through forward computation and selectively integrates it with retained memory via a task-trained recurrent gate. The consolidated memory is then mapped to backbone-specific low-rank adaptation (LoRA) parameters, allowing the encoding capability to transfer when the backbone is replaced. Evaluation across three long-term memory benchmarks and five diverse backbones demonstrates broad generalization with near-constant update cost and memory footprint. With Qwen3-8B on PERMA, RPMem reaches 85.52%, outperforming the strongest parametric and text-based baselines by 5.32 and 12.98 percentage points, respectively. Ablations validate the complementary roles of session compilation and cross-session consolidation, while dynamics analyses reveal that the gate acquires task-specific memory integration strategies. These results establish RPMem as a lifecycle-independent parametric memory framework that maintains evolving cross-session memory that remains reusable across backbone replacements. Our implementation is available at https://github.com/Quark-Medical/rpmem/tree/main.

Comments38 pages, 7 figures. Code: https://github.com/Quark-Medical/rpmem/tree/main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑