arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相关性并非充分性:什么真正弥合了长期记忆问答中的证据缺口

Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

Yufeng Li, Shuxin Li, Zhenhua Xu, Junxian Li, Peng Zeng, Sheng Yao, Changting Lin, Gaolei Li, Ran Bi, Meng Han

arXiv 2610.09348首次发表:更新:

发表机构

East China Normal University; Nanyang Technological University; Zhejiang University; Shanghai Jiao Tong University; University of Southern California; Northeastern University(华东师范大学; 南洋理工大学; 浙江大学; 上海交通大学; 南加州大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长期记忆问答中检索证据相关但不足的问题,提出预算受限的扁平重建方法,通过形式概念分析分解问题并迭代补全证据,显著提升答案质量与证据覆盖率。

AI 中文摘要

与用户跨多个会话交互的LLM智能体所积累的历史记录会超出其上下文窗口,因此它们将过去的交互存储在外部记忆中,并从一小部分检索到的记录中回答每个问题。现有的记忆系统按词法或嵌入相关性对记录进行排序,然而排名靠前的记忆可能各自相关,但共同遗漏了回答问题所需的互补事实,尤其是在多会话和时间性问题中。借鉴法律证据学中相关性与充分性的区分,我们将记忆检索重新定义为构建一个充分的记忆集合。为实现这一观点,我们引入了对检索集合的盲审LLM判断,以及Gold Hit和Turn Hit作为证据覆盖率的代理指标。随后,我们提出了预算受限的扁平重建(BFR)方法,该方法在固定的扁平记忆存储上分两个阶段构建充分集合。具体而言,我们首先应用形式概念分析进行记忆选择(FCA-MS),将问题分解为信息需求,并选择一个紧凑的候选子集以共同覆盖这些需求。然后,我们通过更深入的文本搜索或互补的实体与会话视图反复获取未见过的记录,直到预算耗尽。在LoCoMo和LongMemEval-S上的实验表明,BFR在答案质量和证据覆盖率方面均优于近期智能体记忆系统的同存储适配版本。具体而言,在LongMemEval-S上,它将评判准确率从72.4%提升至82.2%,并将Turn Hit提升至91.4%。

英文摘要

LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions. Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set. To operationalize this view, we introduce a blinded LLM judgment over the retrieved set, together with Gold Hit and Turn Hit as evidence-coverage proxies. We then propose Budgeted Flat Reconstruction (BFR), which builds sufficient sets over a fixed flat memory store in two stages. Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them. Then, we repeatedly acquire unseen records through deeper text search or complementary entity and session views, stopping when the budget is exhausted. Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage. Specifically, on LongMemEval-S it raises judged accuracy from 72.4% to 82.2% and Turn Hit to 91.4%.

Comments25 pages. Code is available at https://github.com/LYF199903/BFR-An-Agent-Memory-Framework-for-Evidence-Reconstruction

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑