可验证记忆:为大语言模型智能体学习结合局部与全局验证器的统一记忆管理
Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
浏览论文内容
中文总结 AI 辅助
提出VerMem框架,用含七个原子操作的统一策略管理LTM等状态,结合局部与全局验证器训练,在多基准测试中性能与效率均优于基线。
中文摘要 AI 辅助
大语言模型(LLM)智能体必须在长时程交互中保留可复用信息、控制有限的活跃上下文,并恢复早期证据。现有方法通常分别优化长时记忆(LTM)和短时记忆(STM),而统一策略往往主要通过轨迹级反馈训练,这为单个记忆决策提供的信用信号较弱。本文提出可验证记忆(VerMem)框架,该框架将长时记忆(LTM)、活跃上下文和片段历史表示为不同状态,并通过一个记忆操作策略对其进行控制。该策略包含七个原子操作,可实现LTM条目的添加、修改或软删除、将LTM检索至活跃上下文、过滤或总结活跃上下文,以及恢复选定的片段片段。VerMem通过监督微调初始化,并采用三阶段强化学习课程进行训练:局部验证器对可执行的记忆转换进行评分,全局验证器在任务完成后评估证据连贯性和终端记忆一致性;这些评分与通过分层信用分配计算的任务、证据召回、效率和约束信号相结合,且验证器仅在训练期间使用。在五个基准测试和两个LLM主干模型上,VerMem在绝大多数报告指标中取得最佳结果,且始终优于强大的记忆基线方法;在三个交互式基准测试的受控在线token预算下,它还在对比方法中实现了最强的效率-性能前沿。代码可在该https URL获取。
英文摘要
Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon interaction. Existing methods commonly optimize long-term memory (LTM) and short-term memory (STM) separately, while unified policies are often trained primarily with trajectory-level feedback, which provides weak credit for individual memory decisions. We present Verifiable Memory (VerMem), a framework that represents LTM, active context, and episodic history as distinct states and controls them with one memory operation policy. Seven atomic operations let the policy add, revise, or soft-delete LTM entries; retrieve LTM into the active context; filter or summarize the active context; and restore selected episodic fragments. VerMem is initialized by supervised fine-tuning and trained with a three-stage reinforcement-learning curriculum. The local verifier scores executable memory transitions, and a global verifier assesses evidence coherence and terminal-memory consistency after task completion. These scores are combined with programmatically computed task, evidence-recall, efficiency, and constraint signals through hierarchical credit assignment. The verifiers are used only during training. Across five benchmarks and two LLM backbones, VerMem achieves the best result on the vast majority of reported metrics and consistently outperforms strong memory baselines. Under controlled online-token budgets on three interactive benchmarks, it also achieves the strongest efficiency--performance frontier among the compared methods. Code is available at https://github.com/Sun-SYSU-24/VerMem.