arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mnemon:原始记录、快速判断、缓慢思考

Mnemon: Raw Records, Fast Judgments, Slow Thoughts

Guangren Wang

arXiv 2609.36059首次发表:更新:

AI 中文总结

Mnemon通过系统1快速判断与系统2缓慢思考的划分,利用原始记录和决策模型实现高效长期记忆,在多个基准上取得领先性能并降低成本。

AI 中文摘要

长期记忆使大语言模型助手能够利用其无法再次阅读的历史,大多数记忆系统在写入时将对话重写为事实、图或类型化记忆。我们认为,记忆的工作如同思考一样,分为两个系统。大部分是快速的系统1工作:对记录进行许多小而独立的“是/否”判断,例如某条记录是否需要或是否已过时,决策模型在三分之一秒内能做出数十个这样的判断。只有少量是缓慢的系统2工作:编写少量搜索查询、命名回复所需的内容并撰写答案,大语言模型擅长但速度较慢。我们提出了Mnemon,一个基于这种划分的记忆代理。它将对话保留为原始的、带日期的记录;一个LLM(系统2)规划对这些记录的搜索,一个决策模型Jev(系统1)判断搜索结果,而带有明确预算的规则将这些判断转化为一个小型视图,供未修改的应答模型使用。后台过程将每条记录整合到主题时间线、价值历史和与记录关联的常设指令中,使得关于整个对话的问题能够触及自身搜索遗漏的证据。由于在写入记录时不对其做任何决定,同一代理可以读取任何返回带日期记录的存储库。使用gpt-4.1-mini进行回答,如在14个系统的公开重新评估中,Mnemon在LoCoMo上得分91.7%,为其中最高,在LongMemEval-S上得分83.8%,每个问题仅使用不到4k个token的上下文,且在LoCoMo上具有最低的有效成本指数。使用推理模型回答时,它在LoCoMo上达到92.2%,在LongMemEval-S上达到94.4%,后者与已发表的最佳结果持平。在BEAM上,从100K到10M token的历史,每个问题的成本仅增长1.11倍。在相同记录上,Jev区分金证据的能力优于两个LLM,且速度快3-11倍。

英文摘要

Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer, which an LLM does well but slowly. We present Mnemon, a memory agent built on this division. It keeps conversations as raw, dated records; an LLM (System 2) plans searches over them, a decision model, Jev (System 1), judges what the searches return, and rules with explicit budgets turn the judgments into a small View for an unchanged answering model. A background pass consolidates each record once into topic timelines, value histories and standing instructions linked to the records, so that questions about a whole conversation reach evidence their own searches miss. Because nothing is decided about a record when it is written, the same agent can read any store that returns dated records. With gpt-4.1-mini answering, as in a public re-evaluation of 14 systems, Mnemon scores 91.7% on LoCoMo, the highest among them, and 83.8% on LongMemEval-S, from under 4k tokens of context per question, with the lowest effective cost index on LoCoMo. With a reasoning model answering, it reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S, the latter on par with the best published results. From 100K to 10M tokens of history on BEAM, its cost per question grows by a factor of 1.11. On the same records, Jev separates gold evidence better than two LLMs and is 3-11 times faster.

Comments16 pages, 3 figures, 4 tables. Code, prompts and run records: https://github.com/Grivn/mnemon-memory-agent

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑