arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择性遗忘:面向长期大语言模型智能体的基于图的记忆框架

Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents

Theo Rusu, Sourena Khanzadeh, Manar Alalfi

arXiv 2608.28978首次发表:更新:

发表机构

Toronto Metropolitan University; The Creative School; Flybits(多伦多都会大学; 创意学院; 弗莱比特公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估了基于图的长期LLM智能体记忆框架的假设,发现其在LongMemEval上未优于扁平向量基线,但遗忘模块能高效剪枝节点且性能损失小,结果受提取器和单一基准限制。

AI 中文摘要

已有研究提出将知识图谱作为长期智能体记忆的结构化替代方案,以替代扁平的检索增强生成,其假设是将对话表示为实体和关系可提升召回率。我们直接评估该假设。我们的框架将每轮对话提取为带类型的节点和带属性的边,从两段子图中回答问题,并定期对在近期性、访问频率、度中心性和年龄的加权组合上得分较低的节点进行剪枝。在LongMemEval基准上,当匹配的候选生成预算为5个检索根时,图结构并未优于扁平向量基线:token F1为0.417,而基线为0.468;对500个问题进行的配对自助抽样显示,差值Δ=-0.050(95%置信区间[-0.085, -0.016])。在需要回忆特定先前助手轮次的问题上,差距最大,此时判断的正确率从0.911降至0.607,这表明将一轮分解为实体会丢弃这些问题所依赖的表面形式。遗忘模块的表现更出色:将其应用于包含27021个节点的持久图时,它移除了9.8%的节点和9.5%的存储字节;token F1保持不变(+0.001,95%置信区间[-0.015, +0.016]),判断的正确率下降1.6个百分点,95%置信区间将任何损失限制在3.8个百分点以内([-0.038, +0.006])。由于我们的提取器是在一个基准上评估的小型单一模型,这些结果表征的是该基于提取的流程,而非一般的图结构记忆。代码:this https URL

英文摘要

Knowledge graphs have been proposed as a structured alternative to flat retrieval-augmented generation for long-term agent memory, on the assumption that representing conversations as entities and relations improves recall. We evaluate that assumption directly. Our framework extracts each conversational turn into typed nodes and attributed edges, answers questions from a two-hop subgraph, and periodically prunes nodes that score low on a weighted combination of recency, access frequency, degree centrality, and age. On LongMemEval, the graph does not outperform a flat vector baseline at a matched candidate-generation budget of five retrieval roots: token F1 is $0.417$ against $0.468$, and a paired bootstrap over 500 questions gives $Δ= -0.050$ (95\% CI $[-0.085, -0.016]$). The gap is widest on questions that require recalling a specific prior assistant turn, where judged correctness falls from $0.911$ to $0.607$, suggesting that decomposing a turn into entities discards the surface form these questions depend on. The forgetting module is more successful. Applied once to a persistent 27{,}021-node graph, it removes 9.8\% of nodes and 9.5\% of stored bytes; token F1 is unchanged ($+0.001$, 95\% CI $[-0.015, +0.016]$) and judged correctness falls by $1.6$ points, with the 95\% interval bounding any loss at $3.8$ points ($[-0.038, +0.006]$). Because our extractor is a single small model evaluated on one benchmark, these results characterise this extraction-based pipeline rather than graph-structured memory in general. Code: https://github.com/skhanzad/Selective-Amnesia

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑