发表机构
Zero Labs(零实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Agent Zero Memory系统,采用三个并行记忆架构,通过溯源规则杜绝幻觉,在LongMemEval和LoCoMo基准上取得最优性能,且成本-效率表现优异。
AI 中文摘要
大语言模型(LLM)智能体需要对用户或组织所说及存储的所有内容具备持久、可信的记忆能力,但多数记忆系统采用单一组织架构(事实存储、向量索引或知识图谱),并继承了该架构的固有缺陷。我们提出Agent Zero Memory,这是一种溯源感知的长期记忆系统,它将用户的对话、文件及关联源提炼为三个并行记忆系统,每个系统捕获同一历史的不同维度:一是将时间和变化内容作为核心要素的事件式记忆事件时间线,二是跨会话关联人物与项目的关联实体-事件知识图谱,三是包含持久事实、带引用锁定的语义精选分层文档式记忆(HDM)。检索流程依次经过意图门(确保独立检索轮次无额外延迟)、源路由器及三个并行的智能体搜索(每个系统对应一个,均为在智能体控制的过滤器下基于混合搜索(嵌入+词汇)的工具使用循环);其基于事实、带引用的答案被整合为一个带有单一置信度的答案。我们明确了阅读规则:所有学习项均为携带来源、时间戳及证据指针的溯源项,所有答案均在引用锁定下读取,仅可引用读者实际打开的证据;该系统从结构上杜绝了生成幻觉,且在无法确定时会弃权(不执行)而非猜测。在两个公开基准上,该系统达到了新的最优性能:在LongMemEval上准确率为95.60%,在LoCoMo上为93.60%,较此前最强系统分别提升0.73和1.10个百分点。针对8种主干LLM的对照研究明确了准确率-成本-延迟的权衡边界:准确率仅波动3.4个百分点,而单查询成本差异约30倍,在单查询成本低至20倍时仍接近最优质量,这体现了记忆驱动而非模型驱动的质量特性。
英文摘要
Large language model (LLM) agents need durable, faithful memory of everything a user or organization has said and stored, yet most memory systems commit to a single organizing structure (a fact store, a vector index, or a knowledge graph) and inherit its blind spots. We present Agent Zero Memory, a provenance-aware long-term memory system that distils a user's conversations, files, and connected sources into three parallel memory systems, each capturing a different facet of the same history: an episodic Memory Events timeline that makes when and what changed first-class, an associative entity-event knowledge graph that links people and projects across sessions, and a semantic, curated, citation-locked Hierarchical Documentary Memory (HDM) of durable facts. A retrieval turn runs an intent gate (so self-contained turns add no latency), a source router, and three concurrent agentic searches, one per system, each a tool-using loop over hybrid (embedding + lexical) search under agent-controlled filters; their grounded, cited answers are integrated into one answer with a single confidence. We formalize the reading discipline: every learned item is a provenanced item carrying its origin, timestamp, and evidence pointer, and every answer is read under a citation lock, so it may cite only evidence its reader actually opened; fabrication is structurally excluded and the system abstains rather than guesses. On two public benchmarks the system sets a new state of the art: 95.60% on LongMemEval and 93.60% on LoCoMo, improving over the strongest prior systems by +0.73 and +1.10 points. A controlled study across eight backbone LLMs characterizes the accuracy-cost-latency frontier: accuracy varies by only 3.4 points while per-query cost varies by ~30x, with near-state-of-the-art quality at up to 20x lower cost per query, the signature of memory-driven, rather than model-driven, quality.
Comments10 pages, 3 figures