发表机构
University of Arkansas at Little Rock(阿肯色大学小石城分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长期记忆智能体检索到不兼容信息的问题,提出检索可接纳性验证框架,通过状态分类和暴露追踪,在公共基准上提升锚点召回率并减少评估开销。
AI 中文摘要
长期记忆智能体可能检索到与当前请求不兼容的相关信息,因为这些信息属于其他主体、违反策略或反映不兼容的生命周期状态。召回率和最终答案准确率无法揭示这一问题:一条路径可能因缺少所需证据而看似安全,而正确答案可能源于不兼容的提示暴露。我们提出一个检索可接纳性验证框架,该框架为每个记忆-查询对分配三种状态之一(可接纳、不可接纳或未解决),在匹配所需证据召回率下比较路径并为未解决情况设定界限,同时通过提示暴露追踪记忆ID并将暴露与目标级披露相关联。我们在独立的非合并群体上评估其各阶段。对来自两个公共长期记忆基准RHELM和MemOps的冻结排名进行事后top-20重分析,覆盖3,767个查询。所有发布的锚点位于可信查询命名空间内;在命名空间内分数不变的情况下,命名空间外过滤无法降低其排名。Top-20锚点召回率从0.432增加到0.533,80%召回可行性从0.237增加到0.311,精确相似度评估减少98.3%。在冻结的72案例开发诊断中,发布元数据参考保留了所需证据,而两个仅文本验证器在1%所需锚点错误拒绝限制下均未检测到违规。在1,523个配对基准原生案例中,命名空间路由与三位读者判断准确率提升0.053-0.068相关;召回率也发生变化,因此该比较为观察性。在16个受控暴露场景中,四位读者特定的95%置信区间中仅有一个排除零,针对相关不可接纳字面披露(+0.156,95% CI [0.031, 0.312])。结果支持对候选支持、可接纳性、提示暴露和答案披露进行单独验证。
英文摘要
Long-term-memory agents can retrieve relevant information that is inadmissible for the current request because it belongs to another principal, violates policy, or reflects an incompatible lifecycle state. Recall and final-answer accuracy do not reveal this: a route can appear safe by missing required evidence, while a correct answer may follow inadmissible prompt exposure. We introduce a retrieval-admissibility verification framework that assigns each memory-query pair one of three statuses (admissible, inadmissible, or unresolved), compares routes at matched required-evidence recall with bounds for unresolved cases, and tracks memory IDs through prompt exposure while linking exposure to target-level disclosure. We evaluate its stages on separate, non-pooled populations. A post-hoc top-20 reanalysis of frozen rankings from two public long-term-memory benchmarks, RHELM and MemOps, covers 3,767 queries. All released anchors lie within trusted query namespaces; with within-namespace scores unchanged, off-namespace filtering cannot lower their ranks. Top-20 anchor recall increases from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations decrease by 98.3%. In a frozen 72-case development diagnostic, a released-metadata reference preserves required evidence, whereas neither text-only verifier detects violations under the 1% required-anchor false-denial limit. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of 0.053-0.068 across three readers; recall also changes, so this comparison is observational. In 16 controlled exposure scenarios, only one of four reader-specific 95% confidence intervals excludes zero for relevant-inadmissible literal disclosure (+0.156, 95% CI [0.031, 0.312]). Results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.
Comments26 pages. Accepted at the NeurIPS 2026 Workshop "Who Verifies the Agents? Toward Reliable Agent Development". Code: https://github.com/ziwang11112/right-memory-wrong-context