发表机构
School of Data Science and Engineering, East China Normal University(数据科学与工程学院,华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型智能体长期记忆中原始对话历史的问题,提出LazyMem方法,将记忆构建推迟到查询时,通过轻量级模型处理候选池,经特定训练方式,在基准测试中取得高准确率,减少记忆令牌数和延迟。
AI 中文摘要
长期记忆使大语言模型智能体能够复用过去的交互,但原始对话历史冗长且信息稀疏。广泛检索可提高证据覆盖率,但会用噪声淹没下游推理;写入时压缩可减少噪声,但会不可逆转地丢弃未来查询可能需要的细节。我们引入LazyMem,通过将所有记忆构建推迟到查询时来避开这一困境。一个轻量级的4B模型在重叠的并行窗口中处理检索到的候选池,仅选择性地保留和压缩与查询相关的内容。该模型通过监督微调,然后是基于组的强化学习进行训练,使用格式门控复合奖励,该奖励结合了基于规则的衡量选择准确性的动作信号和由大语言模型判断的衡量源可信度和查询效用的质量信号。在LongMemEval基准测试中,LazyMem-4B仅用213个记忆令牌就达到了0.85的大语言模型判断准确率,比仅检索少68.7倍,并且在没有目标域训练的情况下推广到LoCoMo(0.68),同时降低了先前查询时基线的平均延迟。32B变体达到0.93,在聚合密集型问题类型上超过了 oracle-context参考。这项工作的相关代码可在这个https URL上公开获取。
英文摘要
Long-term memory enables LLM agents to leverage past interactions, but dialogue histories quickly exceed the context window, forcing agents to retrieve relevant subsets at query time. Because useful evidence is sparse and scattered across verbose conversations, retrieval faces a fundamental tension: broadening recall improves coverage but floods downstream reasoning with noise, while compressing memories at write time eases retrieval but irreversibly discards details that future queries may need. We introduce LazyMem, which resolves this tension by deferring all memory construction to query time. Given a retrieved candidate pool, a lightweight model processes it in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained with supervised fine-tuning followed by reinforcement learning, using a reward that jointly encourages the identification of relevant messages and the generation of compressions that are faithful to the source and useful for answering the query. On LongMemEval, LazyMem-4B achieves an LLM-judge accuracy of 0.85, outperforming the strongest non-oracle baseline while using only 213 answer-context memory tokens, 21.0 times fewer than the baseline. It further generalizes to LoCoMo without target-domain training and reduces mean latency relative to the prior query-time baseline. Code is available at https://github.com/allacnobug/LazyMem.
Comments29 pages, under review