LycheeMemory V2:基于语义片段级整合的LLM智能体高效长期记忆
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
浏览论文内容
中文总结 AI 辅助
LycheeMemory V2用语义片段级整合替代轮次级整合,降低LLM编码成本,在LoCoMo、LongMemEval-S上实现SOTA性能,构建token用量较A-Mem大幅减少且未增加查询开销。
中文摘要 AI 辅助
长周期LLM智能体必须保留过往交互信息以支持未来任务。现有记忆系统通常依赖即时整合,每次交互后调用LLM来提取、总结或更新记忆,这种设计会使记忆构建成本随对话增长而不断升高。粗粒度总结可降低构建成本,但存在丢失细粒度上下文证据的风险;而扩大检索上下文或多跳LLM推理则会将开销转移到查询阶段。我们提出LycheeMemory V2,一种高效长期记忆框架,它用语义片段级整合替代了轮次级整合。LycheeMemory不会对每一次交互进行整合,而是将多轮对话批量处理为片段,并将每个已完成的片段编码为与上下文无关的类型化记忆记录。与固定窗口批量处理相比,片段级批量处理降低了LLM编码频率,而语义边界检测有助于保留连贯的事件级和时间证据。生成的记录通过轻量级结构化索引组织,用于按查询需求检索证据。使用GPT-4.1-Mini开展的实验表明,LycheeMemory实现了最先进的性能,在LoCoMo上达到89.22%,在LongMemEval-S上达到92.20%;与A-Mem相比,它在LoCoMo上减少了86.0%的构建token,在LongMemEval-S上减少了75.9%的构建token,且未增加查询阶段的token使用量。更广泛地说,我们的结果表明,长期智能体记忆的准确率-成本权衡不仅取决于保留的信息,还取决于整合的粒度。
英文摘要
Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory systems typically rely on eager consolidation, invoking LLMs after each interaction to extract, summarize, or update memories. This design makes memory construction increasingly costly as conversations grow. Coarse summarization can reduce construction cost but risks discarding fine-grained contextual evidence, whereas larger retrieval contexts or multi-hop LLM reasoning shift the overhead to query time. We present LycheeMemory V2, an efficient long-term memory framework that replaces turn-level consolidation with semantic segment-level consolidation. Instead of consolidating every interaction, LycheeMemory batches multiple exchanges into segments and encodes each finalized segment into context-independent typed memory records. Segment-level batching lowers LLM encoding frequency, while semantic boundary detection helps preserve coherent event-level and temporal evidence compared with fixed-window batching. The resulting records are organized with lightweight structured indexes for query-planned evidence retrieval. Experiments using GPT-4.1-Mini show that LycheeMemory achieves state-of-the-art performance, reaching 89.22% on LoCoMo and 92.20% on LongMemEval-S. Compared with A-Mem, it reduces construction tokens by 86.0% on LoCoMo and 75.9% on LongMemEval-S without increasing query-time token usage. More broadly, our results suggest that the accuracy--cost trade-off of long-term agent memory depends not only on what information is retained, but also on the granularity at which it is consolidated.
发表机构
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。