发表机构
University of Southern Queensland; Southern University of Science and Technology; Jiangsu University; The University of Sydney; Uploading Inc.(南昆士兰大学; 南方科技大学; 江苏大学; 悉尼大学; Uploading 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PACMI通过溯源图与四状态有效性格实现级联内存失效,提升长期LLM智能体处理过时内存的准确性,在诊断基准上取得最优结果。
AI 中文摘要
LLM智能体在执行长期任务时依赖长期内存来保留和重用信息。现有方法在处理因新观察或领域证据到达而过时的内存方面支持有限。这些过时的内存可能在语义上仍然相关,继续影响依赖记录,并作为历史证据保留价值。这需要两种能力:依赖追踪以识别下游影响,以及历史保留以保留有用的过去记录。我们提出了溯源感知级联内存失效(PACMI),一个将内存和新证据表示为具有类型化依赖边的溯源图的框架。PACMI将记录分配到四状态有效性格中,将有效性变化传播到依赖内存,并使用结果状态进行检索和过时前提检测。我们还引入了一个包含五个领域100个案例和300个查询的诊断基准。评估分别考察节点级、上下文级和答案级性能。PACMI在该基准上实现了最高的最终答案准确率,其与最强基线的配对差异在精确McNemar检验下显著。前提检查器在受控查询分布上实现了完美的精确率、召回率和F1分数。级联传播主要改善了内存状态正确性:移除它会使最终答案错误从3增加到11,但配对差异未达到0.05显著性阈值。代码和数据将公开发布。
英文摘要
LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memories may remain semantically relevant, continue to affect dependent records, and retain value as historical evidence. This calls for two capabilities: dependency tracking to identify downstream effects and historical preservation to retain useful past records. We propose Provenance-Aware Cascading Memory Invalidation (PACMI), a framework that represents memories and new evidence in a provenance graph with typed dependency edges. PACMI assigns records to a four-state validity lattice, propagates validity changes to dependent memories, and uses the resulting states for retrieval and stale-premise detection. We also introduce a diagnostic benchmark with 100 cases and 300 queries across five domains. The evaluation separates node, context-, and answer-level performance. PACMI achieves the highest final-answer accuracy on this benchmark, and its paired difference from the strongest baseline is significant under an exact McNemar test. The premise checker achieves perfect precision, recall, and F 1 on the controlled query distribution. Cascading propagation primarily improves memorystate correctness: removing it increases final-answer errors from 3 to 11, but the paired difference does not reach the 0.05 significance threshold. Code and data will be made publicly available.