arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemRetriever:学习从长期记忆中搜索、反思与检索

MemRetriever: Learning to Search, Reflect, and Retrieve from Long-Term Memory

Ruiyang Jiang, Chunyu Li, Zhiyu Li

arXiv 2609.11951首次发表:更新:

发表机构

MemTensor (Shanghai) Technology(明策(上海)科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MemRetriever智能体检索模型,将长期记忆访问视为多步搜索过程,通过并行/串行搜索与反思去噪动态决策,并用组相对策略优化训练,在多数据集上优于静态检索基线。

AI 中文摘要

长期记忆使个性化智能体成为可能,但其价值取决于在正确的时间检索到正确的证据。大多数记忆系统使用静态的 top-k 检索:它们发出一个查询,返回固定数量的记忆,并直接将其传递给下游模型。这种方法可能会遗漏分布在多个会话中的证据,引入不相关的内容,并浪费上下文,尤其是对于多跳、时间性和知识更新类问题。我们提出了 MemRetriever,一种智能体检索模型,将记忆访问视为一个多步骤的搜索过程。在每一步中,MemRetriever 基于当前证据进行推理,并选择并行搜索以进行广泛探索、串行搜索以进行针对性补全,或反思与去噪以进行过滤和证据评估。当保留的证据足以用于下游回答时,它便停止。我们构建了 ReAct 风格的搜索-记忆轨迹,用于监督式冷启动训练,并进一步使用组相对策略优化(Group Relative Policy Optimization)来优化模型。奖励设计鼓励证据覆盖、噪声减少、答案充分性和高效终止。在 LOCOMO、LongMemEval、HotpotQA、MuSiQue 和 2WikiMultiHopQA 上的实验表明,与静态检索和仅监督式基线相比,该方法取得了持续改进。MemRetriever-4B-RL 在相同流程下,在 LongMemEval 的主要检索指标上还优于 DeepSeek-v4-Flash,并在 MuSiQue 上取得了所比较方法中最强的结果。由于其决策逻辑独立于存储后端,MemRetriever 还可以在外部知识库和向量数据库上运行。这些结果表明,一个规划、搜索、过滤证据并决定何时停止的中间决策层,可以改进长期记忆检索和知识密集型问答。

英文摘要

Long-term memory enables personalized agents, but its value depends on retrieving the right evidence at the right time. Most memory systems use static top-k retrieval: they issue one query, return a fixed number of memories, and pass them directly to a downstream model. This approach can miss evidence distributed across sessions, introduce irrelevant content, and waste context, especially for multi-hop, temporal, and knowledge-update questions. We present MemRetriever, an agentic retrieval model that treats memory access as a multi-step search process. At each step, MemRetriever reasons over the current evidence and selects parallel search for broad exploration, serial search for targeted completion, or reflection and denoising for filtering and evidence assessment. It stops when the retained evidence is sufficient for downstream answering. We construct ReAct-style search-memory trajectories for supervised warm-start training and further optimize the model with Group Relative Policy Optimization. The reward design encourages evidence coverage, noise reduction, answer sufficiency, and efficient termination. Experiments on LOCOMO, LongMemEval, HotpotQA, MuSiQue, and 2WikiMultiHopQA show consistent improvements over static retrieval and supervised-only baselines. MemRetriever-4B-RL also outperforms DeepSeek-v4-Flash on the main LongMemEval retrieval metrics under the same pipeline and achieves the strongest results among the compared methods on MuSiQue. Because its decision logic is independent of the storage backend, MemRetriever can also operate over external knowledge bases and vector databases. These results show that an intermediate decision layer that plans, searches, filters evidence, and determines when to stop can improve both long-term memory retrieval and knowledge-intensive question answering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑