arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10413cs.AI

幸运回忆:面向LLM持久连贯性的本体驱动记忆生命周期管理

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Ansuman Mullick, Eray Tüzün

首次发表
浏览论文内容

中文总结 AI 辅助

提出本体驱动的记忆生命周期管理方法,按行为类型分类个人事实并应用差异化策略,在多个基准上提升检索正确率并大幅降低虚构。

中文摘要 AI 辅助

当前的LLM记忆系统将所有个人事实等同对待,导致存储无界增长而检索精度下降。核心挑战在于生命周期管理:根据每个事实的行为类型,哪些记忆应持久保留、哪些应被替换、以及以何种速率进行。幸运回忆(Fortunate Recall, FR)是一个可组合的策略层,将个人事实分类为10+1行为本体,并应用类别特定的生命周期策略(差异化时间衰减、槽位键替代、事件时间有效性和类别感知检索路由),作为LLM提取元数据上的确定性函数。FR-Bank是我们与基础设施无关的实现,在LifecycleBench(一个包含516个问题的时间消歧新基准)上达到76.9%的通过率,领先于Mem0、A-MEM、Memory-R1和MemoryOS(61%至70.5%),并在规范Wu等人评判协议下于完整LongMemEval-S上达到75.2%,表明生命周期策略对标准检索无可测量成本。预注册消融实验定位了增益来源:用三个通用生命周期原语替换类型化层后,正确性统计上无显著变化(-1.7个百分点,95%置信区间[-6.0, +2.7]),因此通用生命周期元数据承载正确性优势,而行为本体承载校准能力,将下游虚构减半(12.0%对24.2%,p<0.001)。端到端地,FR-Bank在已回答查询上将虚构从Mem0的45.1%降至22.4%,在所有查询上从32.2%降至13.0%,同时正确回答更多查询(31.2%对18.6%);排名在开放权重Kimi K2.5上复现。该分解迁移至独立构建的BEAM基准:在280个问题上正确率46.8%,优于Mem0的32.9%,本体收益集中于矛盾消解,并在约七个策略簇处饱和。本体、基准和代码已发布。

英文摘要

Current LLM memory systems treat all personal facts identically, so stores grow without bound while retrieval precision degrades. The core challenge is lifecycle management: which memories should persist, which should be replaced, and at what rate, conditioned on the behavioral type of each fact. Fortunate Recall (FR) is a composable policy layer that classifies personal facts into a 10+1 behavioral ontology and applies category-specific lifecycle policies (differential temporal decay, slot-key supersession, event-time validity, and category-aware retrieval routing) as deterministic functions over LLM-extracted metadata. FR-Bank, our infrastructure-independent implementation, reaches a 76.9% pass rate on LifecycleBench, a new 516-question temporal-disambiguation benchmark, ahead of Mem0, A-MEM, Memory-R1, and MemoryOS (61% to 70.5%), and 75.2% on the full LongMemEval-S under the canonical Wu et al. judge protocol, so lifecycle policies impose no measurable cost on standard retrieval. A pre-registered ablation locates the gains: replacing the typed layer with three generic lifecycle primitives leaves correctness statistically unchanged (-1.7pp, 95% CI [-6.0, +2.7]), so the generic lifecycle metadata carries the correctness advantage, while the behavioral ontology carries calibration, halving downstream confabulation (12.0% vs 24.2%, p<0.001). End-to-end, FR-Bank cuts confabulation from Mem0's 45.1% to 22.4% over answered queries and from 32.2% to 13.0% over all queries while answering more of them correctly (31.2% vs 18.6%); the ranking replicates on the open-weight Kimi K2.5. The decomposition transfers to BEAM, an independently built benchmark: 46.8% correct vs Mem0's 32.9% over 280 questions, with the ontology's benefit concentrated in contradiction resolution and saturating near seven policy clusters. The ontology, benchmark, and code are released.

发表机构

  • Bilkent University(比尔肯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑