arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于大语言模型智能体的准确高效长期记忆

Accurate and Efficient Long-Term Memory for LLM Agents

Zicheng Zhao, Xinyang Guo, Luyao Lv, Menghan Wang, Ming Li, Shuaicheng Li

arXiv 2607.16211首次发表:更新:

发表机构

Central China Research Institute for AI Technology; Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science, BNU-HKBU United International College, Zhuhai; City University of Hong Kong(华中人工智能技术研究院; 北京师范大学-香港浸会大学联合国际学院数据科学跨学科研究与应用广东省重点实验室; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型智能体长时记忆问题,提出MOSAIC框架,通过实体类型图存储、哈希加速双路径检索和主动冲突检测,提升了准确性和效率,在多项测试中取得更好结果,如准确率提高、冲突检测能力增强且检索延迟降低。

AI 中文摘要

增强持久记忆的大语言模型智能体可召回过去交互,但现有系统有局限:扁平无结构存储丢失多跳和时间推理所需关系上下文,依赖昂贵的基于大语言模型的分类使其不适用于对延迟敏感的部署,且无验证机制会积累矛盾。我们提出MOSAIC,一种结构化、冲突感知的大语言模型智能体长时记忆框架,更准确高效。它引入三项关键能力:实体类型图存储、哈希加速双路径检索和主动冲突检测。在多项测试中,MOSAIC取得了更好的结果,如在LoCoMo上准确率达89.35%,在HaluMem上有最佳提取F1值,能检测66%的注入事实冲突,同时哈希加速检索使平均搜索延迟低至每秒0.58个问题。

英文摘要

LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructured storage loses relational context needed for multi-hop and temporal reasoning, and reliance on expensive LLM-based classification makes them impractical for latency-sensitive deployment. Without mechanisms to validate new information against stored knowledge, these systems silently accumulate contradictions. We present MOSAIC (Memory-Organized Structured Agent for Information Collection), a structured, conflict-aware long-term memory framework for LLM agents that is substantially more accurate and efficient. MOSAIC introduces three key capabilities: (1) entity-typed graph storage with semantic classification preserving relational structure across events, personas, and relationships, enabling multi-hop and temporal reasoning over conversation history; (2) hash-accelerated dual-path retrieval replacing LLM-based classification with locality-sensitive hashing, achieving near-instantaneous lookup with negligible accuracy loss; and (3) active conflict detection at save time that cross-references new information against existing graph neighbors, triggering updates or deletions for contradictory entries. Evaluated on LoCoMo (long-conversation QA), HaluMem, and a novel clinical-guideline error compounding test, MOSAIC achieves 89.35% accuracy on LoCoMo (+27.21 pp over the best baseline), best HaluMem-Medium extraction F1(86.77%) and HaluMem-Long extraction F1 (85.84%), best QA correctness on both Medium and Long (73.10%, 70.75%), and detects 66% of injected factual conflicts-4.7 times higher than the best baseline (14%)-while hash-accelerated retrieval keeps average search latency at 0.58 s per question.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑