FluctlightDB:面向AI智能体的数据内存模型
FluctlightDB: A Memory Model of Data for AI Agents
AI总结:
本文提出面向AI智能体的独立数据内存模型,实现嵌入式引擎FluctlightDB,在多个记忆召回基准上取得优异结果,填补了数据栈的一层空白。
AI中文摘要:
五十年来,数据系统始终围绕两个核心问题构建:关系模型用于查询哪些记录匹配给定谓词,向量模型用于查询哪些向量与查询向量距离最近。但这两类模型均未针对智能体在长会话中基于线索驱动、结合来源权重的记忆召回场景设计。本文提出将智能体的长期记忆视为一种独立的数据模型,具备专属的写入语义(编码、分离、整合、来源追溯)与读取语义(基于线索在关联记忆图中激活),并实现了名为 FluctlightDB 的嵌入式引擎,通过 experience() 与 activate() 接口实现上述约定。本文谨慎论证该方案,未断言其创新性超越 Mem0、Zep 或 HippoRAG 风格的记忆层,仅聚焦于这些上层方案之下的嵌入式引擎约定。实验结果显示:在 LoCoMo(官方证据召回指标,含10个对话、1982个黄金片段)上,CHORUS在2026年7月的内部复现运行中召回率达99.0%;在 LongMemEval-S(500个问题,官方 session_recall@8)上,本文的检索框架得分97.6%(488/500),结合阅读器/判断器栈的端到端问答得分97.4%(487/500)——这些层采用的协议与本文仅作参考引用的供应商排行榜数据不同;在 BEIR SciFact(共享 MiniLM 嵌入、相同框架、开启 Recall Fabric)上,CHORUS/PRISM 在 nDCG@10(0.646 vs. 0.645)与 Recall@10(0.792 vs. 0.783)指标上优于 Chroma。本文还报告了作者设计的小型回归测试集 FAMB(复述测试 n=10,其他子测试 n=1),其宏观得分100%,为内部验证结果,非同行基准。用户可通过 pip install "fluctlightdb[native]" 及极简的 connect() -> experience() -> activate() 脚本(编译 wheel,非仅源码)在一分钟内验证该引擎。框架与冻结 JSON 采用 MIT 许可。本文未提出新的神经科学理论或新的 Transformer 模型,仅提出数据栈中缺失的一层,并发布可供他人复现与验证的引擎。
英文摘要:
For fifty years, data systems have answered two questions. The relational model asked which records match a predicate; the vector model asked which vectors lie nearest a query. Neither was built for cue-driven, provenance-weighted recall across long sessions. We propose treating long-term agent memory as a distinct data model -- with its own write semantics (encoding, separation, consolidation, provenance) and read semantics (cue-driven activation across a linked memory graph) -- and present FluctlightDB, an embedded engine that implements this contract via experience() and activate(). We do not claim novelty over Mem0, Zep, or HippoRAG-style memory layers, only an embedded engine contract beneath them. Numbers are typed by metric; retrieval is not generation. On LoCoMo (official evidence-recall; 10 conversations, 1,982 gold spans), our native-Rust CHORUS stack reaches 96.8% at k=150 as raw evidence recall with no neighbor expansion, on an internally reproduced July 2026 run; at k=5 it still yields 72.6%, while end-to-end QA over date-stamped context reaches 85% at k=15 (retrieval-bound). On LongMemEval-S (500 questions), official session_recall@8 is 97.6% (488/500) and end-to-end QA with our reader/judge stack is 97.4% (487/500) -- different protocols from vendor leaderboard figures we cite for context only. On BEIR SciFact (shared MiniLM embeddings, same harness), CHORUS/PRISM edges Chroma on nDCG@10 (0.646 vs. 0.645) and Recall@10 (0.792 vs. 0.783). A graded provenance-conflict suite (n=50) scores 18% top-1 when all pairs share one brain versus 100% under per-case isolation (ceiling, not deployment evidence). Engine, harnesses and frozen JSON are MIT; pip install "fluctlightdb[native]" re-runs the published numbers. We claim no new neuroscience and no new transformer: a missing layer of the data stack, released for others to re-run and contest.