arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习检索长期记忆问答中缺失的证据

Learning to Retrieve Missing Evidence for Long-Term Memory QA

Yi-Xuan Deng, Yi Zhang, Wei Liu, Chao Xue, Shuojin Yang

arXiv 2609.37443首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长期记忆问答中证据缺失问题,提出MERA方法,通过强化学习训练轻量级规划器,基于已验证证据引导检索,显著提升答案准确率和证据召回率。

AI 中文摘要

长期记忆使语言模型能够在未来对话中使用过去的交互。然而,回答一个问题所需的证据可能分散在相隔较远的轮次中,而问题本身又省略了定位这些证据所需的线索。检索到的事实可以揭示这些线索,从而激励基于已发现证据的检索决策。我们引入了MERA(缺失证据检索增强),它将全局可搜索记忆与特定于问题的证据状态分离。已验证的证据指导后续检索,同时不限制对全局记忆的访问。我们通过强化学习训练一个轻量级规划器,奖励能够恢复先前缺失证据的查询。MERA在Qwen3-30B和GPT-4o-mini骨干模型上实现了强大的答案准确性。使用Qwen3-30B进行证据处理和答案生成时,经过训练的0.6B规划器在LoCoMo上达到77.40%的准确率,在LongMemEval-S上达到71.29%,分别超过未经过检索接地训练的30B规划器4.10%和3.96%。在LoCoMo上,后续检索轮次将累积证据召回率从55.5%提升至80.5%。

英文摘要

Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.

Comments22pages,6figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑