发表机构
Salesforce AI Research(赛富时人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过冻结的最后写入解析器执行 MemoryAgentBench 冲突解决基准,发现官方指标下规则可答对 80.25% 问题,但 67 个不可达项目暴露了可达性分割和解析器残留,强调逐项分割而非总体分数才是解读关键。
AI 中文摘要
MemoryAgentBench 的冲突解决子集被解读为衡量“选择性遗忘”。我们执行基准自身的规则——关于某个事实的最新陈述获胜——作为一个零学习解析器,冻结在四个事实列表中的一个上。在官方指标下,该规则回答了 80.25% 的问题(在三个保留列表上为 74.5%)。其余部分中,有 67 个项目具有已发布的黄金答案,但最后写入图无法达到,而覆盖的陈述可以达到(例如,“印度的首都是新德里”被“印度的首都是格罗塞托”取代;黄金答案是新德里);此类项目在 262K 的多跳问题中占三分之一。两个长上下文模型和我们预先注册的基准 BM25 智能体的近似重新实现(每个项目仅保留一次运行和结果)在规则解决的问题上得分分别为 84.7%、82.6% 和 41.6%,而在那 67 个项目上得分分别为 10.4%、11.9% 和 6.0%。失败是一个可达性分割加上一个小的解析器范围残留;逐项分割,而非总体,是此处分数可被解读的单位。
英文摘要
MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.
CommentsAccepted at the IAB Workshop (Interpreting Agent Behavior) at NeurIPS 2026 (non-archival). 19 pages