arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TAGGRAPH:用于智能体持久历史图检索的标签增强图

TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories

Yu-Shu Chen, Yu-Jung Liang, Pengtao Xie

arXiv 2609.38353首次发表:更新:

发表机构

University of California, San Diego(加州大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TAGGRAPH,通过受控评估框架比较图检索与词法基线在智能体长期记忆中的表现,发现检索策略需与记忆设置联合评估,并强调词汇规范化与提取质量的重要性。

AI 中文摘要

长期记忆使LLM智能体能够回忆过去的交互并在会话间保持一致,但记忆系统难以比较,因为它们通常在表示、索引、检索和评估方面各不相同。我们提出一个基于共享5W风格对话记忆的受控评估框架。局部化图配置遍历一个公共基础图;AdaptiveGraph添加时间顺序边和个性化PageRank扩散。我们还评估了在相同提取笔记上的BM25以及OpenClaw作为原始输入的外部参考。检索排名因记忆设置而异。在LongMemEval-S上,AdaptiveGraph是最强的图配置,MRR为0.844,但BM25达到0.867,OpenClaw达到0.880。在ATANT Core上,局部化图遍历优于扩散和BM25,而BM25在压力轮次中领先。在测试范围内减小LongMemEval-S并不能重现ATANT的扩散惩罚,但测试的最小存储仍大于ATANT Core,因此不能排除存储大小的影响。该惩罚在宽松的内容匹配标准下仍然存在。词汇规范化和提取质量显著影响图检索,而缺失的提取标签在排名前五的未命中中很常见。因此,检索策略应与记忆设置联合评估,并与强词法基线进行比较。

英文摘要

Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized graph configurations traverse a common base graph; AdaptiveGraph adds chronological edges and Personalized PageRank diffusion. We also evaluate BM25 over the same extracted notes and OpenClaw as a raw-input external reference. Retrieval rankings vary across memory settings. On LongMemEval-S, AdaptiveGraph is the strongest graph configuration at 0.844 MRR, but BM25 reaches 0.867 and OpenClaw 0.880. On ATANT Core, localized graph traversal outperforms diffusion and BM25, whereas BM25 leads the stress rounds. Reducing LongMemEval-S within the tested range does not reproduce the ATANT diffusion penalty, but the smallest tested store remains larger than ATANT Core, so store size cannot be ruled out. The penalty also persists under a permissive content-match criterion. Vocabulary normalization and extraction quality substantially affect graph retrieval, and missing extraction tags are common among top-five misses. Retrieval strategies should therefore be evaluated jointly with the memory setting and against strong lexical baselines.

CommentsAn earlier version was accepted at the COLM 2026 Workshop on Lifelong Learning Agents (LLA)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑