arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

牢记于心:评估智能体记忆中的隐式关联盲点

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Ruizhe Li, Mingxuan Du, Benfeng Xu, Zhendong Mao

arXiv 2607.24368首次发表:更新:

发表机构

University of Science and Technology of China; Metastone Technology(中国科学技术大学; 元石科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究智能体记忆中的隐式关联盲点问题,引入InMind基准测试,通过配对控制区分三种解释,实验发现将记忆置于上下文时主干回答间接查询比例高,检索相同记忆时其他系统召回率低,诊断探针可缩小差距,指出路由是关键问题。

AI 中文摘要

长期记忆系统将用户所说内容存储在外部存储中,并在相关查询到来时进行检索。该接口基于一个很少被提及但很自然的假设:所需的记忆与需要它的查询相似。然而,世界知识打破了这一假设。例如,对坚果过敏应改变对含有杏仁粉成分的马卡龙请求的回答,但这两个文本没有检索器能看到的线索。我们将这种失败模式称为隐式关联盲点,并引入了InMind,这是一个经过专家验证的包含125个任务的基准测试,涵盖十个生活领域,其中113个任务基于可引用的公共来源。其配对控制区分了现有评估中混淆的三种解释:事实从未存储、模型缺乏桥接知识或事实已存储但未浮出水面。结果很清晰。当将决定性记忆置于上下文中时,主干能回答84.0%的间接查询;而当必须检索相同记忆时,六个向量、图形和智能体记忆系统最多只能达到14.4%,尽管它们按需回忆相同事实的比例高达100%。维度增加八倍的嵌入提高了每个系统对答案盲目标的召回率,但差距基本保持不变。一个在查询到来之前保持记忆可见的最小诊断探针缩小了大部分差距,将故障定位在查询条件接口本身,并指出路由(决定哪些事实必须保持可见)是InMind旨在评估的开放问题。

英文摘要

Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑