智能体应该记住什么?在有限记忆评估中区分保留与检索
What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation
浏览论文内容
中文总结 AI 辅助
本研究通过流式回忆基准区分保留与选择,发现固定访问时查询感知选择提升回忆15.5个百分点,并建议有限记忆评估应分别报告保留与选择。
中文摘要 AI 辅助
一个持久运行的智能体必须同时决定在信息到达时保留什么,以及在查询出现时呈现什么,然而记忆评估可能通过比较在保留和选择两方面都不同的方法而混淆这些决策。我们构建了一个流式回忆基准,交叉组合保留规则与选择规则,并在相同的300个种子化情节上评估每种条件。在保持访问方式固定的情况下,查询感知选择将所需事实的回忆率提高了15.5个百分点(95%置信区间:12.8至18.2),而同时改变历史访问方式的混合比较则报告了68.7个百分点的优势,其中53.2个百分点可归因于访问方式。在有限保留条件下,查询感知选择、密集选择和oracle选择均达到保留上限,且有限近因条件中观察到的全部319次失败均由驱逐而非排序错误导致。当目标足够久远地退入过去时,回忆率降至0%。在SQuAD上重复评估保留了保留上限,同时表明密集检索在自然文本上可优于词法检索。这些结果表明,有限记忆评估应保持访问方式固定,并分别报告保留与选择的结果。
英文摘要
A persistent agent must decide both what to retain as information arrives and what to surface once a query appears, yet memory evaluations can confound these decisions by comparing methods that differ in both retention and selection. We build a streaming-recall benchmark crossing retention and selection rules and evaluate every condition on the same 300 seeded episodes. Holding access fixed, query-aware selection improves required-fact recall by 15.5 percentage points (95% CI: 12.8 to 18.2), whereas a mixed comparison that also changes history access reports a 68.7-point advantage, of which 53.2 points are attributable to access. Under bounded retention, query-aware, dense, and oracle selection reach the retention ceiling, and all 319 observed failures in the bounded recency condition are caused by eviction rather than ranking errors. Recall falls to 0% as targets recede sufficiently far into the past. Repeating the evaluation on SQuAD preserves the retention ceiling while showing that dense retrieval can outperform lexical retrieval on natural text. These results show that bounded-memory evaluations should hold access fixed and report retention and selection separately.
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。