CMI-Mem:通过CMI增强的强化学习实现可泛化的长期内存管理
CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
研究针对现有内存管理器模型局限,提出基于强化学习的CMI-Mem模型,结合下游问答正确性与内在条件互信息CMI作为混合奖励,可评估新输入信息,补充而非取代问答基础,实现可泛化的长期内存管理。
中文摘要 AI 辅助
内存管理器模型在智能体系统中至关重要。现有方法主要依赖大语言模型判断的合成问答对,使内存评估依赖采样查询和下游阅读器。为解决此局限,我们提出CMI-Mem,一种基于强化学习的轻量级内存管理器模型,具有结合下游问答正确性和内在条件互信息(CMI)的混合奖励。CMI评估新对话输入相对于当前内存状态所贡献的信息,而无需依赖采样问答查询,从而补充而非取代问答基础。我们的代码可在[this https URL]获取,CMI-Mem-4B模型检查点可在[this https URL]获取。
英文摘要
Memory Manager models are pivotal in agent systems. Existing reinforcement-learning methods commonly use LLM-judged synthetic question-answer (QA) pairs: this provides useful downstream task grounding, but values memory through a sampled query distribution and a fixed reader. We propose CMI-Mem, a lightweight RL memory manager with a hybrid reward. Its extrinsic QA term measures end-task correctness, while its intrinsic Conditional Mutual Information (CMI) term evaluates the information contributed by new conversational inputs relative to the current memory state without conditioning on a sampled QA query. The two signals are complementary: QA anchors task utility, whereas CMI provides per-operation supervision for relevant, non-redundant memory construction. Experiments demonstrate improved transfer across memory-use scenarios, together with more efficient training and inference from the per-operation CMI signal. Our codes are available at: https://github.com/Wyb0627/CMIMem , and the CMI-Mem-4B model checkpoint is available at: https://www.modelscope.cn/models/wyb0627/CMIMem-4B
发表机构
- Alibaba Group(阿里巴巴集团)
- HKUST(香港科技大学)
- PolyU(香港理工大学)
- HKUST(GZ)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。