arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用记忆:记忆智能体中记忆载体的整体评估

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

Wei-Chieh Huang, Weizhi Zhang, Yuchen Wu, Yankai Chen, Eric Hanchen Jiang, Wooseong Yang, Yiwei Yang, Henry Peng Zou, Hanrong Zhang, Ying Nian Wu, Haolun Wu, Kai-Wei Chang, Philip S. Yu, Xue Liu, Aylin Caliskan

arXiv 2608.15008首次发表:更新:

发表机构

University of Illinois Chicago; University of Washington; McGill University; MBZUAI; University of California, Los Angeles(伊利诺伊大学芝加哥分校; 华盛顿大学; 麦吉尔大学; 穆罕默德·本·扎耶德人工智能大学; 加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究评估了记忆智能体的多种记忆载体,发现无通用最优载体,不同载体适配不同场景,为自适应智能体记忆系统设计提供了实证指导。

AI 中文摘要

记忆正成为长视野大语言模型(LLM)智能体的核心基础设施,但现有评估对不同运行场景下应使用哪种记忆载体(即记忆被表示和存储的底层介质)提供的指导有限。我们对记忆增强型智能体的记忆载体开展了受控的harness评估,涵盖密集索引、稀疏索引、文本记录、结构化存储、分层存储、基于精调的记忆、参数更新以及与激活兼容的上下文机制。在三个骨干模型和四个基准套件(涵盖以用户为中心的问答和以智能体为中心的决策)上,我们在统一的harness下测量了26项性能和效率指标。结果表明,没有任何一种载体能始终占据优势:广泛的检索有益于长上下文事实问答,但过度检索会通过将注意力从对决策关键的上下文转移开而损害序列决策。可扩展性引入了另一个路由维度,因为在中等历史长度下表现良好的载体在更长视野下可能变得成本高昂或脆弱。这些发现促使载体路由成为自适应智能体记忆系统的必要组成部分,并为设计高效、可靠且适配场景的LLM智能体长期记忆提供了实证指导。代码将在论文接收后提供。

英文摘要

Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used under different operating regimes. We present a controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms. Across three backbone models and four benchmark suites spanning user-centric question answering and agent-centric decision-making, we instrument 26 performance and efficiency metrics under a unified harness. Our results show that no single substrate consistently dominates: broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context. Scalability introduces a further routing axis, as substrates that perform well at moderate history lengths can become costly or brittle at longer horizons. These findings motivate substrate routing as a necessary component of adaptive agent memory systems and provide empirical guidance for designing efficient, reliable, and regime-aware long-term memory for LLM agents. Code will be made available upon acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑