arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34438cs.CLcs.AI

通过提问来记忆:面向LLM智能体的检索诱导记忆演化

Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents

Wanqi Zhou, Jiawei Lu, Yang Wang, Zhaolong Xing, Zhen Chen, Ai Han, Haoyue Shi

AI总结:

提出RIME框架,通过检索诱导的证据中心记忆整合替代整体压缩,并在推理时补充源对话检索,在LoCoMo上以更少令牌实现最优性能。

AI中文摘要:

长期记忆对于语言智能体在长时间、多会话交互中维持连贯且有效的行为至关重要。现有的记忆系统主要在读时使用检索,而写时的记忆形成仍依赖于直接提取或压缩。然而,当未来的信息需求未知时,一次性压缩整个交互可能会忽略那些日后可能重要的局部细节。为此,我们提出了RIME,一种检索诱导记忆框架,将记忆构建从整体压缩转向以证据为中心的整合。RIME使用通用的自问问题来检索聚焦的对话证据,并将记忆形成建立在检索到的证据和相关历史记忆之上,这些内容被联合协调成一个带有时间和来源信息的不断演化的记忆库。在推理时,压缩记忆作为主要而非唯一的证据来源:当它无法支持答案时,RIME会检索相关的源对话及其局部上下文,以恢复记忆形成过程中遗漏的信息,而无需进行全历史处理。在LoCoMo上使用Qwen3-235B-A22B和GPT-5.6 Sol进行的大量实验表明,RIME在所有三个质量指标上始终优于对比方法,同时所需的查询时LLM令牌数大幅减少。

英文摘要:

Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.

↑