arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20553cs.AIcs.CL

CMI-Mem:通过CMI增强的强化学习实现可泛化的长期内存管理

CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning

Yubo Wang, Qiuyu Zhao, Zenghui Sun, Shichao Dong, Jinsong Lan, Xiaoyong Zhu, Haoyang Li, Bo Zheng, Lei Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现有内存管理器模型局限,提出基于强化学习的CMI-Mem模型,结合下游问答正确性与内在条件互信息CMI作为混合奖励,可评估新输入信息,补充而非取代问答基础,实现可泛化的长期内存管理。

中文摘要 AI 辅助

内存管理器模型在智能体系统中至关重要。现有方法主要依赖大语言模型判断的合成问答对,使内存评估依赖采样查询和下游阅读器。为解决此局限,我们提出CMI-Mem,一种基于强化学习的轻量级内存管理器模型,具有结合下游问答正确性和内在条件互信息(CMI)的混合奖励。CMI评估新对话输入相对于当前内存状态所贡献的信息,而无需依赖采样问答查询,从而补充而非取代问答基础。我们的代码可在[this https URL]获取,CMI-Mem-4B模型检查点可在[this https URL]获取。

英文摘要

Memory Manager models are pivotal in agent systems. Existing reinforcement-learning methods commonly use LLM-judged synthetic question-answer (QA) pairs: this provides useful downstream task grounding, but values memory through a sampled query distribution and a fixed reader. We propose CMI-Mem, a lightweight RL memory manager with a hybrid reward. Its extrinsic QA term measures end-task correctness, while its intrinsic Conditional Mutual Information (CMI) term evaluates the information contributed by new conversational inputs relative to the current memory state without conditioning on a sampled QA query. The two signals are complementary: QA anchors task utility, whereas CMI provides per-operation supervision for relevant, non-redundant memory construction. Experiments demonstrate improved transfer across memory-use scenarios, together with more efficient training and inference from the per-operation CMI signal. Our codes are available at: https://github.com/Wyb0627/CMIMem , and the CMI-Mem-4B model checkpoint is available at: https://www.modelscope.cn/models/wyb0627/CMIMem-4B

发表机构

  • Alibaba Group(阿里巴巴集团)
  • HKUST(香港科技大学)
  • PolyU(香港理工大学)
  • HKUST(GZ)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

↑