arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CMT-RAG:面向多轮多跳检索增强生成(RAG)的互补记忆轨迹

CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG

Lang Zhou, Yingjian Chen, Shuxuan Li, Kun-Yu Lin, Zhilin Zhao

arXiv 2607.26470首次发表:更新:

AI 中文总结

该研究针对现有RAG系统难以恢复多轮对话中先前推理与证据的问题,提出CMT-RAG框架,引入MuMu-QA基准,实验显示其在答案准确性上优于五类RAG基线。

AI 中文摘要

多轮信息检索对话既需要多跳推理,也需要跨轮次的长程依赖跟踪。然而,现有RAG系统通常将对话记忆表示为原始对话历史、重写查询或非结构化摘要,难以恢复后续查询所需的特定先前推理步骤和证据。我们的核心洞见是通过将对话上下文表示为子问题级推理轨迹,使对话记忆与检索实现对齐。基于此,我们引入MuMu-QA——一个带有显式跨轮子问题依赖标注的多轮多跳RAG基准,以及面向该场景的互补记忆框架CMT-RAG。在每一轮,CMT-RAG会使用状态空间轨迹生成器,其循环状态作为运行时记忆,整合近期对话上下文并将当前查询分解为结构化轨迹草稿,这些草稿包含面向检索的子问题及对先前轨迹的依赖关系。随后,它利用检索到的证据对这些草稿进行落地,并将它们存储为会话级有向无环图(DAG)中的持久记忆轨迹,使后续轮次能够高效恢复相关的先前推理与证据。在MuMu-QA及语料库级RAG基准上的实验表明,CMT-RAG在答案准确性上始终优于五类RAG基线。

英文摘要

Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG systems typically represent conversational memory as raw dialogue history, rewritten queries, or unstructured summaries, making it difficult to recover the specific prior reasoning steps and evidence required for follow-up queries. Our key insight is to align conversational memory with retrieval by representing dialogue context as sub-question-level reasoning traces. Building on this insight, we introduce MuMu-QA, a benchmark for multi-turn multi-hop RAG with explicit cross-turn sub-question dependency annotations, and CMT-RAG, a complementary memory framework for this setting. At each turn, CMT-RAG employs a state-space trace generator, whose recurrent state serves as runtime memory, to incorporate recent conversational context and decompose the current query into structured trace drafts containing retrieval-oriented sub-questions and dependencies on earlier traces. It then grounds these drafts with retrieved evidence and stores them as persistent memory traces in a session-level DAG, enabling future turns to efficiently recover relevant prior reasoning and evidence. Experiments on MuMu-QA and corpus-level RAG benchmarks show that CMT-RAG consistently outperforms five categories of RAG baselines in answer accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑