arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21962cs.CL

真实优先:智能体记忆的纵向评估工具及记忆架构排名中的任期交叉

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

Quentin Spencer

AI总结:

研究大型语言模型智能体记忆基准问题,提出颠倒流程的评估方法,通过合成虚构语料库嵌入新特征,对五种记忆架构基准测试,发现排名随历史长度反转,分层架构表现最佳,发布开源库Veracium。

AI中文摘要:

大型语言模型智能体记忆基准通常先生成对话,再提取答案键,存在标签错误和污染问题,且多衡量短交互历史。本文颠倒流程:先由种子生命脚本采样器在文本生成前发出带有效区间、波动类别和源渠道的事实;再由语言模型渲染器根据事实清单编写聊天和邮件;保真度验证器确认每个植入事实;问题由脚本机械实例化,黄金答案通过构建有效且单独验证可回答性。合成虚构语料库嵌入了所调查基准中没有的特征。通过对五种记忆架构与无记忆控制进行基准测试发现,后端排名随历史长度反转,编写阶段质量与下游质量强相关,注入抗性跟踪来源边界是否在表示中幸存。分层架构在两种情况下表现最佳,并作为开源库Veracium发布,附带语料库生成器和工具。

英文摘要:

Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contamination problems -- and they overwhelmingly measure short interaction histories. We invert the pipeline: a seeded life-script sampler emits facts with validity intervals, volatility classes, and source channels before any text exists; an LLM renderer writes chat and email from per-event fact manifests; a fidelity verifier confirms every planted fact; and questions are instantiated mechanically from the script, so gold answers are script-valid by construction and separately validated for answerability. The synthetic, fictionalized corpus (~380 questions, 15 types) embeds features absent from the benchmarks we survey: per-fact validity intervals, sent/received trust distinctions, injection probes in a benign harness, and as-of-date question sets. Benchmarking five memory architectures against a no-memory control (fixed answerer, versioned LLM judge, three replicates, two horizons), we find backend rankings invert with history length: the budgeted curated-map memory that leads at three weeks loses recall of evicted content by nine weeks (96% to 72%) while a provenance-typed graph rises to 90%; the inversion is positive for all six users under complete cross-family re-judging (exact p=0.031). A full-rendered-history baseline ties or exceeds the best memory system at the short horizon but shows no judge-independent advantage at nine weeks, at about twice the read cost. Write-stage quality strongly correlates with downstream quality (weakly-written facts fail 24% vs 2%), and injection resistance tracked whether provenance boundaries survive representation. A layered architecture performs best among the memory systems in both regimes (96.8% short-horizon) and is released as Veracium, an open-source library, with the corpus generator and harness.

补充信息

↑