发表机构
Davidson College(戴维森学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对对话式AI智能体跨会话持久内存缺失的问题,提出原生图双时态内存存储方案,经LongMemEval基准测试,在知识更新等任务上取得良好效果,为纯检索的局限性提供了改进方向。
AI 中文摘要
对话式AI智能体普遍缺乏跨会话的持久内存。常见解决方案如将完整聊天历史注入上下文窗口,或委托第三方内存服务,要么耗尽模型的上下文预算,要么通过用户无法控制的基础设施传输个人数据。本文描述一种可避免上述问题的内存存储:智能体本地的Neo4j属性图,辅以HNSW向量索引和完整双时态数据模型。每条内存存储为不可变身份节点,链接到带两个闭开时间区间的版本化内容节点:有效时间(事实在世界中为真的时间)和事务时间(数据库记录该事实的时间)。该设计支持时点语义检索,无需物理覆盖历史。写入时,通过1024维嵌入的余弦相似度自动维护相关内存间的语义边。我们在LongMemEval上评估该系统,这是一个包含6种问题类型、共500道题的基准,旨在测试长期记忆能力。在60道抽样问题中,当前状态语义搜索路径的整体R@10为46.7%,在知识更新问题上升至80%;时间旅行路径在知识更新问题上R@10为80%,但在时间推理问题上召回率从50%降至37.5%,这是后过滤稀释导致的,直接指明了具体的设计改进方向。我们讨论这些结果揭示的纯检索在不同问题类型上的局限性,以及每种失败模式对未来工作的启示。
英文摘要
Conversational AI agents commonly lack persistent memory across sessions. The obvious fixes like injecting full chat histories into the context window, or delegating to a third-party memory service, either exhaust the model's context budget or send personal data through infrastructure the user does not control. We describe a memory store that avoids both problems: an agent-local Neo4j property graph augmented with HNSW vector indexes and a full bitemporal data model. Each memory is stored as an immutable identity node linked to versioned content nodes carrying two closed-open time intervals: valid time (when the fact was true in the world) and transaction time (when the database recorded it). This design supports point-in-time semantic retrieval without physically overwriting history. Semantic edges between related memories are maintained automatically at write time using cosine similarity over 1024-dimensional embeddings. We evaluate the system on LongMemEval, a 500-question benchmark spanning six question types designed to stress long-term memory. Across 60 sampled questions, the current-state semantic search path achieves 46.7% R@10 overall, rising to 80% on knowledge-update questions. The time-travel path yields 80% R@10 on knowledge-update but decreases recall on temporal-reasoning questions (50% to 37.5%), a consequence of post-filter dilution that points directly to a concrete design improvement. We discuss what these results reveal about the limits of pure retrieval for different question types and what each failure mode suggests for future work.