AI 中文总结
针对LLM智能体对话的时效性、来源追溯等问题,提出基于证据的时序图记忆SodaMem,在LongMemEval-S上以Flash级花费达到高准确率,性能优于部分高成本系统。
AI 中文摘要
在数周对话中为用户提供协助的大语言模型(LLM)智能体,必须记住当前为真的内容,而非仅仅是曾说过的内容。平面RAG日记和Markdown日志优化了“针检索”(needle retrieval),但在时效性、来源追溯和有序时序推理方面表现不足(Maharana等人,2024;Wu等人,2024;Packer等人,2023;Chhikara等人,2025)。本文提出SodaMem,一种基于证据的时序图记忆,其具备三个核心特性:(i)提取带有强制来源范围的类型化事实事件(FactEvents);(ii)在混合词汇-密集索引下,持久化提及时间、发生时间和有效性,并设置SUPERSEDES/ CONTRADICTS/ UPDATES边;(iii)通过规划器-读取器循环进行响应,该循环在生成最终回复前收集可引用的证据。在LongMemEval-S数据集上,采用store-of-record配置时,使用deepseek-v4-flash模型的准确率达到92.8%(500个样本中正确464个,为3次运行中的最佳结果),平均每问题成本为0.00161美元(约18.3k token;中位数成本0.00111美元,约14.6k token)。本文将带有可估算API成本的公开系统编译为成本表和成本-准确率图;根据这些估算,SodaMem在Flash级别的花费下处于准确率前沿附近,且在多个更高成本、更低准确率的系统中表现严格占优。准确率使用与读取器、评估器相同的Flash模型(自评估);成本不包含摄取/评估环节,跨系统比较为编译估算而非单一框架。代码可在指定的两个网址获取。
英文摘要
Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat RAG diaries and Markdown logs optimize needle retrieval but under-serve currency, provenance, and ordered temporal reasoning (Maharana et al. 2024; Wu et al. 2024; Packer et al. 2023; Chhikara et al. 2025). We present SodaMem, an evidence-grounded temporal graph memory that (i) extracts typed FactEvents with mandatory provenance spans, (ii) persists mention time, occurrence time, and validity with SUPERSEDES/CONTRADICTS/UPDATES edges under hybrid lexical-dense indexing, and (iii) answers via a planner-reader loop that gathers citable evidence before composing a final response. On LongMemEval-S, our store-of-record configuration reaches 92.8% accuracy (464/500; best of N=3) at mean $0.00161/question (approximately 18.3k tokens; median $0.00111 / approximately 14.6k) with deepseek-v4-flash. We compile public systems with estimable API cost into a cost table and cost-accuracy map; under these estimates SodaMem sits near the accuracy frontier at Flash-tier spend and strictly dominates several higher-cost, lower-accuracy points. Accuracy uses the same Flash model as reader and judge (self-grading); costs exclude ingest/judge and cross-system comparisons are compiled estimates rather than a single-harness bake-off.Our code is available at https://github.com/SodaMem/SodaMem