发表机构
Megagon Labs(美格根实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体记忆驱逐导致的遗忘,提出恢复反事实审计方法,区分可恢复与不可逆错误,发现多数错误不可逆,且预算-准确率结果受检索机制影响。
AI 中文摘要
智能体记忆系统在历史记录超过固定令牌预算时,必须丢弃已存储的信息。现有的预算-准确率前沿量化了由此产生的准确率损失,但并未区分由驱逐(eviction)导致的不可逆损失与可恢复的检索失败。我们引入了恢复反事实(restore counterfactual),这是一种针对每个问题的配对干预措施,它将问题的黄金证据(gold evidence)重新置于读取时的上下文中,并重新运行相同的读取器。将正确性的变化与驱逐后证据是否被保留相结合,可将每个可被神谕回答(oracle-answerable)的错误分类为可恢复的、不可逆的或残余的;在残余情况下,即使恢复证据后答案仍然不正确。我们使用GPT-4o-mini作为主要读取器和评判器,GPT-5.4-mini作为稳健性读取器,在LongMemEval-S上,以三种预算和两种检索机制下,评估了FIFO、随机、冗余感知和LLM重要性驱逐策略。在80k令牌预算下的top-k检索中,对于FIFO、随机和冗余感知驱逐,被恢复纠正的错误中不可逆的比例为0.67-0.73,而LLM重要性驱逐为0.60。在8k令牌预算下,所有四种策略的该比例均达到1.00。可恢复错误在80k令牌预算下的top-k检索中出现,但在强制黄金证据注入(forced-gold injection)下按构造不存在,因此除非报告检索机制,否则预算-准确率结果不可直接比较。一项探索性的匹配准确率分析在1.2-6个百分点的分辨率下,未检测到准确率匹配的策略对之间不可逆率的差异。同一分析检测到了故意破坏性对照(deliberately destructive control)的差异。据我们所知,这是首次针对标准对话基准上的外部智能体记忆存储,进行逐项、逐问题的驱逐恢复反事实审计。
英文摘要
Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing budget-accuracy frontiers quantify the resulting loss in accuracy, but do not distinguish irreversible losses caused by eviction from recoverable retrieval failures. We introduce the restore counterfactual, a per-question paired intervention that reinstates the question's gold evidence in the read-time context and reruns the same reader. Combining the change in correctness with whether the evidence was retained after eviction classifies each oracle-answerable error as recoverable, irreversible, or residual; in the residual case, the answer remains incorrect after restoration. We evaluate FIFO, random, redundancy-aware, and LLM-importance eviction on LongMemEval-S at three budgets and under two retrieval regimes, using GPT-4o-mini as the primary reader and judge and GPT-5.4-mini as a robustness reader. Under top-k retrieval at an 80k-token budget, the irreversible share among errors corrected by restoration is 0.67-0.73 for FIFO, random, and redundancy-aware eviction, compared with 0.60 for LLM-importance. At 8k tokens, it reaches 1.00 for all four policies. Recoverable errors occur under top-k retrieval at 80k tokens but are absent under forced-gold injection by construction, so budget-accuracy results are not directly comparable unless the retrieval regime is reported. An exploratory matched-accuracy analysis detects no difference in irreversible rate among accuracy-matched policy pairs at a resolution of 1.2-6 percentage points. The same analysis detects the deliberately destructive control. To our knowledge, this is the first per-item, per-question restore-counterfactual audit of eviction for external agent-memory stores on a standard conversational benchmark.