发表机构
Illinois Institute of Technology(伊利诺伊理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对记忆增强语言模型中未检索记忆效用不可识别的问题,提出因果记忆策略(CMP),通过干预检索并采用自归一化逆倾向加权估计记忆效用,显著提升记忆区分度,并揭示识别效用不足以支持保留决策。
AI 中文摘要
记忆增强的大语言模型必须决定保留哪些记忆,最近的系统通过估计每条记忆对任务表现的影响来做出决策。然而,这些估计完全依赖于被检索到的记忆。当一条记忆从未被检索到时,存储层面的干预会产生相同的结果,导致其效用无法被识别。这是检索层面的积极性违反,而仅检查记忆操作的诊断方法无法察觉。我们引入了因果记忆策略(CMP),这是一种因果框架,通过干预检索本身来恢复可识别性,为以已知倾向采样的记忆保留固定数量的上下文槽位。CMP通过平衡分配设计下的自归一化逆倾向加权来估计记忆效用。我们证明了通过检索对记忆效用的因果分解、估计量的无偏性和精确方差,以及在不可逆操作下的最优决策规则。实验上,在LongMemEval上54%的必需记忆和LoCoMo上67%的必需记忆存在识别失败,且该失败在已部署的记忆系统中持续存在。CMP将必需与非必需记忆的区分度从0.54 AUC提升到0.66 AUC。最后,我们表明仅识别出的记忆效用不足以做出保留决策:每条查询的效用在其估计的查询上达到0.78 AUC,但保留策略可用的任何聚合都无法预测记忆在未见查询上的价值。代码可在以下网址获取:此https URL。
英文摘要
Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory's value on unseen queries. Code is available at: https://anonymous.4open.science/r/cmp-release-D0C3/.