arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37930cs.CLcs.LG

学习记住什么:长视野反事实记忆优化

Learning What to Remember: Long-horizon Counterfactual Memory Optimization

Jiaming Tang, Mingyan Liu, Armin Sarabi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出记忆增益策略优化(MGPO),通过边际贡献信用分配解决长交互中记忆重写的延迟效用问题,在文档级信息抽取中提升性能并减少近80%记忆长度,支持跨领域复用。

中文摘要 AI 辅助

持久化的文本记忆使语言模型能够在长时间交互中携带信息,但学习记住什么本质上是一个信用分配问题。一次记忆重写可能要在许多步骤之后才变得有用,而观察到的效用中很大一部分可能继承自重写之前已存储的信息。我们引入了记忆增益策略优化(MGPO),该方法通过将每次记忆重写归因于其对当前和未来下游效用的边际贡献,来隔离其增量价值。这将延迟的记忆效用转化为直接的学习信号,用于优化哪些信息应该持久化。我们在文档级信息抽取上研究了MGPO,其中结构化监督使得单个记忆更新的效果可以直接测量。MGPO在提高抽取性能的同时,相对于优化前的初始记忆策略,平均记忆长度减少了近80%。学习到的记忆策略还支持跨领域和跨下游模型的复用与迁移,无需进一步训练。这些结果表明,有效的记忆学习不仅依赖于保留有用信息,还依赖于识别哪些记忆更新能产生持久的增量价值。

英文摘要

Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

↑