发表机构
UNIST(蔚山科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多跳问答中检索与生成阶段实体信息丢失的问题,提出双实体恢复RAG(DER-RAG),通过双向查询分解和主语实体前缀,无需图构建或微调,在三个基准上达到或超越强基线。
AI 中文摘要
检索增强的多跳问答(QA)将查询分解为子问题,并将语料库分解为更小的检索单元(如句子)。这两种分解方式都改善了流水线,但我们表明,它们共享同一个弱点,即实体信息的丢失,并且这种丢失会在两个不同的节点破坏流水线。第一个节点是检索,子问题丢失了在上一跳中解析出的实体,使得检索器没有可匹配的内容。第二个节点较难察觉,因为检索看起来仍然成功。一旦段落被拆分为句子,孤立的句子就丢失了为其代词提供指代背景的上下文,因此即使正确的句子在手,大语言模型(LLM)也无法判断该句子涉及哪个实体。我们将第二个节点单独界定为一种独特的失败模式,称为“生成中丢失”(lost-in-generation),一项受控检索实验表明,即使黄金证据固定在上下文中,这种失败模式也会降低答案质量。随后,我们提出了双实体恢复RAG(DER-RAG),它通过两个轻量级组件,使接地实体从检索到生成始终保持显式:一个双向查询分解,将已解析的实体跨子问题传递;以及一个在生成时附加到每个句子的主语实体前缀。DER-RAG无需构建图、无需修改语料库、无需微调,但在三个多跳问答基准上,它达到或超过了强基线,包括依赖昂贵离线结构的基于图的方法。
英文摘要
Retrieval-augmented multi-hop question answering (QA) decomposes a query into sub-questions and decomposes the corpus into smaller retrieval units such as sentences. Both forms of decomposition improve the pipeline, but we show that both share the same vulnerability, the loss of entity information, and that this loss breaks the pipeline at two separate points. The first point is retrieval, where a sub-question loses the entity resolved at the previous hop, leaving the retriever with nothing to match against. The second point is harder to see, because retrieval still appears to succeed. Once a passage is split into sentences, an isolated sentence loses the context that grounds its pronouns, so even with the correct sentence in hand the LLM cannot tell which entity the sentence is about. We isolate this second point as a distinct failure mode that we call lost-in-generation, and a retrieval-controlled experiment shows that it degrades answers even when the gold evidence is fixed in the context. We then propose Dual Entity Recovery RAG (DER-RAG), which keeps the grounding entity explicit from retrieval through to generation with two lightweight components, a two-way query decomposition that carries the resolved entity across sub-questions and a subject entity prefix attached to each sentence at generation time. DER-RAG needs no graph construction, no corpus modification, and no fine-tuning, yet on three multi-hop QA benchmarks it matches or exceeds strong baselines, including graph-based methods that depend on costly offline structures.
CommentsFindings of the Association for Computational Linguistics: EMNLP 2026