arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

了解存储:在智能体能够读取之前,记忆后端必须写入什么

Knowing the Store: What a Memory Backend Must Write Down Before an Agent Can Read It

Ansuman Mullick, Eray Tüzün

arXiv 2610.04794首次发表:更新:

发表机构

Bilkent University(比尔肯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨记忆后端需暴露何种信息(如生命周期计数和属性列表)使智能体在检索前识别过时记录,通过元记忆监控实验发现,存储写入的内容限制了所有测试模型的表现。

AI 中文摘要

一个具有长期记忆的智能体可以从它不应再使用的记录中作答,例如用户后来取消的计划。我们探究记忆存储必须暴露什么信息,才能使智能体在检索任何内容之前知晓这一点,并将该判断单独评分,作为元记忆监控。读者仅能看到无价值的摘要:按生命周期状态划分的记录计数和属性名称列表。148个问题中的每一个都针对同一存储的、在某一属性上有所不同的三个版本提出,因此措辞无法泄露答案。模型所增加的内容取决于该列表的长度。在基准测试的ground truth产生的短列表上(平均2.6个名称),五个语言模型均未能可靠地超越对名称的余弦相似度查找,且三个较强的模型与之等效(差异在0.05以内)。在固定计数和列表长度的控制条件下,没有模型可靠地优于该查找,也没有模型被证明与之等效。在比我们测试的存储所写入的更长的列表(填充至60个名称)上,查找损失0.18;GPT-5.6 Luna和Sol损失0.06至0.07,并领先0.13至0.14,而其他三个模型与之同步下降。对于Sol,这一领先在与填充至相同长度和计数的同类模型对比时依然保持。当重新措辞以与名称无共享词时,查找损失0.02,一个较强的读者以0.04的微弱优势领先。在短列表上,三个较强的读者仅在摘要计数某个状态而未命名该状态时领先,且一个计数特征在一个问题内缩小了这一领先,尽管在项目汇总时并非如此。我们测试的后端省略或错误陈述了这一信息:在FR-Bank自身的元数据下,GPT-4.1 mini、Haiku 4.5和查找分别从0.73、0.82和0.75降至0.60、0.71和0.59。在这些测试中,存储所写入的内容限制了我们所提供的每个读者。所有存储均为合成的;更好的判断是否能带来更好的答案这一控制,留待后续研究。

英文摘要

An agent with long-term memory can answer from a record it should no longer use, such as a plan the user later cancelled. We ask what a memory store must expose for an agent to know this before retrieving anything, and we score that judgment on its own, as metamemory monitoring. Readers see only a value-free summary: record counts by lifecycle state and a list of attribute names. Each of 148 questions is asked against three versions of one store that differ in one attribute, so wording cannot give the answer away. What a model adds depends on the length of that list. On the short lists the benchmark's ground truth produces (2.6 names on average), none of five language models reliably beats a cosine-similarity lookup over the names, and the three stronger ones are equivalent to it within 0.05. Under a control that fixes counts and list length, none is reliably above it and none is shown equivalent. On lists longer than the stores we tested write, padded to 60 names, the lookup loses 0.18; GPT-5.6 Luna and Sol lose 0.06 to 0.07 and lead it by 0.13 to 0.14, while the other three fall with it. The lead holds for Sol against a sibling padded to the same length and counts. Reworded to share no word with the names, the lookup loses 0.02 and one stronger reader edges 0.04 ahead. On the short lists the three stronger readers lead only where the summary counts a state without naming it, and one count feature closes that lead within a question, though not when items are pooled. The backends we tested omit or misstate this information: under FR-Bank's own metadata, GPT-4.1 mini, Haiku 4.5 and the lookup fall from 0.73, 0.82 and 0.75 to 0.60, 0.71 and 0.59. In these tests, what the store wrote down limited every reader we gave it to. All stores are synthetic; control, whether a better judgment yields a better answer, is left to a later study.

Comments94 pages. Under review at ICLR 2027. Supplement included as ancillary file. Code and data: https://github.com/ansh348/knowing-the-store

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑