探测智能体记忆中的稳定性-可塑性权衡:基于认知实验范式
Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
浏览论文内容
中文总结 AI 辅助
针对智能体记忆评估忽视稳定性-可塑性权衡的问题,提出认知启发的MemProbe框架,通过四个实验范式将总分分解为行为特征,并在56情节套件中揭示相似分数下的不同记忆维护模式。
中文摘要 AI 辅助
智能体记忆系统越来越多地被用于维护长期用户偏好、任务状态和不断演变的事实,但当前的评估往往将记忆行为简化为最终答案的准确性。我们引入了MemProbe,一个受认知科学启发的框架,用于诊断智能体记忆中的稳定性-可塑性权衡。该框架的动机源于认知记忆研究的一个核心见解:记忆是重构性的,并受到干扰、来源可靠性、强化和再激活的影响。MemProbe将此见解转化为四个可复用的实验范式(干扰、错误信息、巩固强度和再巩固窗口),这些范式操纵了记忆应何时被更新、保留或视为不确定。它进一步将正确性分解为行为特征,揭示系统如何更新、保留、归因和时间上组织信息。我们在一个包含56个情节的诊断套件中实例化了这些范式,并在统一协议下评估了六个增量记忆系统。结果表明,具有相似总分数的系统表现出不同的行为特征。MemProbe提供了这样一个诊断视角,将总体性能转化为随时间推移的记忆维护的可解释特征。代码可在该https URL获取。
英文摘要
Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation. MemProbe turns this insight into four reusable experimental paradigms (interference, misinformation, consolidation strength, and reconsolidation window) that manipulate when a memory should be updated, preserved, or treated as uncertain. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information. We instantiate these paradigms in a 56-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles. MemProbe provides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time. Code is available at https://github.com/jq-ding/MemProbe.
发表机构
- UNC-Chapel Hill(北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。