CoEM:利用基于证据的承诺记忆增强长上下文推理
CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory
浏览论文内容
中文总结 AI 辅助
CoEM提出基于证据的承诺记忆机制,通过待处理集合保留源摘录,用学习策略和冻结验证器决定记忆承诺,结合强化学习训练,在长上下文推理中显著提升性能。
中文摘要 AI 辅助
长上下文推理对于复杂且长期的任务至关重要,然而随着上下文长度的增加,大型语言模型(LLMs)的性能会下降。近期的方法通过逐块处理输入,同时在模型上下文中维护有界的文本记忆来应对这一挑战。然而,过早的信息压缩可能会丢弃对后续推理至关重要的关键细节。在本文中,我们引入了基于证据的承诺记忆(CoEM),它学习何时将源证据转换为紧凑的记忆事实。具体来说,在固定的上下文记忆预算下,CoEM将可能有用的源摘录逐字保留在待处理集合中,允许后续上下文在不可逆压缩之前澄清其相关性。随着新上下文的到来,学习到的策略会重新审视每个待处理的摘录,并决定是将其提升为承诺记忆、保留以供进一步考虑,还是丢弃。一个冻结的验证器确保提议的事实只有在得到保留的摘录和当前上下文支持时才会被接受。为了进一步指导有效的记忆管理,我们通过结合细粒度的、步骤级别的证据奖励和最终答案奖励,使用强化学习来训练该策略。大量实验表明,CoEM持续改善长上下文推理。在6,400个文档的长上下文输入上评估时,CoEM在Qwen3.5-9B上比最强的记忆基线高出10.4-11.4个F1分数点。代码仓库:此https URL。
英文摘要
Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: https://github.com/benmagnifico/CoEM.
发表机构
- Korea Advanced Institute of Science & Technology(韩国科学技术院)
- University of Macau(澳门大学)
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。