LeanMem:面向LLM智能体的简单高效长期记忆系统
LeanMem: Simple and Efficient Long-Term Memory for LLM Agents
浏览论文内容
中文总结 AI 辅助
LeanMem是面向LLM智能体的轻量级长期记忆框架,通过差异化处理历史内容、动态分配资源,在两个数据集上提升了记忆基线的准确率,同时降低了成本与延迟。
中文摘要 AI 辅助
长期记忆对于基于大语言模型(LLM)的智能体维持交互并可靠利用远期历史信息至关重要。然而,现有记忆系统通常通过统一的总结与检索流程处理异构对话内容,要么导致过多的令牌消耗,要么造成细粒度证据的不可逆损失。我们认为,应根据历史对话内容的可压缩性、时间动态性和保真度要求对其进行差异化处理。基于这一见解,我们提出了LeanMem,一个轻量级长期记忆框架。LeanMem首先过滤低价值内容,随后根据信息性质将信息片段存储为紧凑的概要记忆、时间结构化的事件记忆或基于源的记录记忆。维护阶段仅选择性更新动态演化的事件记忆,避免对稳定的概要和不可变的记录进行冗余整合。推理阶段,LeanMem根据查询特定的证据需求动态选择记忆类型并分配检索预算,按需组装相关证据。在LoCoMo和LongMemEval-S数据集上,使用GPT-4.1-mini和Qwen3-8B模型时,LeanMem在所有设置下均优于最强的基于记忆的基线,准确率提升最高达15.1个百分点,同时具备最低或接近最低的构建成本、推理令牌数和延迟。代码和数据集包含在补充材料中。
英文摘要
Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typically process heterogeneous dialogue content through a uniform summarization and retrieval pipeline, leading to either excessive token consumption or irreversible loss of fine-grained evidence. We argue that historical dialogue content should be handled differently according to its compressibility, temporal dynamics, and fidelity requirements. Based on this insight, we propose LeanMem, a lightweight long-term memory framework. LeanMem first filters out low-value content, then stores informative segments as compact profile memory, temporally structured event memory, or source-grounded record memory, depending on the nature of the information. During maintenance, only dynamically evolving event memories are selectively updated, avoiding redundant consolidation of stable profiles and immutable records. During inference, LeanMem dynamically selects memory types and allocates retrieval budgets according to query-specific evidence demands, assembling relevant evidence on demand. On LoCoMo and LongMemEval-S with GPT-4.1-mini and Qwen3-8B, LeanMem improves accuracy over the strongest memory-based baseline in every setting, by up to 15.1 points, at the lowest or near-lowest construction cost, inference tokens, and latency. The code and datasets are included in the supplementary materials.