arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10265cs.AI

过时、错配或迟到:个人记忆在生成之前失效之处

Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation

Haonan Deng, Park Sinchaisri

首次发表
浏览论文内容

中文总结 AI 辅助

该研究直接测量语言代理个人记忆在生成前出现的过时、错配和迟到错误,发现时间有效性主要依赖记忆构建,并主张在生成前评估记忆,分离状态有效性、身份解析、弃权和服务延迟。

中文摘要 AI 辅助

语言代理的个人记忆通常通过最终答案是否正确来评判。这种评分掩盖了生成之前出现的错误:记忆块可能包含过时的值、关于错误人的事实,或在服务截止时间前没有有用的事实。我们直接测量这些失败。使用个人事实记忆(PFM)作为参考层,我们发现时间有效性在我们的设置中主要是记忆构建的属性。在受控修订基准上,仅提供每个正确键控槽的活动值消除了观察到的过时暴露;如果没有更新解析,70.3%的提示会暴露一个被取代的值。一旦检索器共享相同的活动存储和参与者信息,参与者感知的BM25在预设的0.02边际内等同于参考排序器。更难的问题是将修订分配给正确的槽位。遗漏的合并使过时的值保持活动,而错误的合并则静默移除当前值;四个LLM键分配器实现了比规则提取器更高的键召回率,但产生更低的干净检索率,并且LongMemEval上的开放域合并召回率从未超过0.062。错误归属在有效性过滤后仍然存在:实体后验减少了受控数据上的同名暴露,但无法区分LoCoMo中同名的说话者。两个冻结的语言模型在生成的文本中重现了提示错误。检索延迟因排序器而异,但在我们的硬件上,提示预填充主导了轮次级别的延迟。这些结果主张在生成之前评估代理记忆,分离存储状态有效性、身份解析、弃权(不执行)和服务延迟。

英文摘要

Personal memory for language agents is usually judged by whether the final an- swer is correct. That score hides errors that arise before generation: the memory block may contain an obsolete value, a fact about the wrong person, or no use- ful fact before the serving deadline. We measure these failures directly. Using Personal Fact Memory (PFM) as a reference layer, we find that temporal validity is primarily a property of memory construction in our setting. On a controlled revision benchmark, serving only the active value of each correctly keyed slot eliminates observed stale exposure; without update resolution, 70.3% of prompts expose a superseded value. Once retrievers share the same active store and par- ticipant information, participant-aware BM25 is equivalent to the reference ranker within a prespecified 0.02 margin. The harder problem is assigning revisions to the right slot. Missed merges leave stale values active, whereas false merges silently remove current values; four LLM key assigners achieve higher key re- call than a rule extractor yet produce lower clean-retrieval rates, and open-domain merge recall on LongMemEval never exceeds 0.062. Misattribution survives va- lidity filtering: an entity posterior reduces same-name exposure on controlled data but cannot distinguish identically named speakers in LoCoMo. Two frozen lan- guage models reproduce prompt errors in generated text. Retrieval latency varies across rankers, but prompt prefill dominates turn-level latency on our hardware. These results argue for evaluating agent memory before generation, separating stored-state validity, identity resolution, abstention, and serving latency.

发表机构

  • University of California, Berkeley(加州大学伯克利分校)
  • Haas School of Business, University of California, Berkeley(加州大学伯克利分校哈斯商学院)

机构由 AI 辅助整理,请以论文原文为准。

↑