arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17043cs.CLcs.AIcs.IR

诊断多跳问答中的事实锚定差距

Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering

Kevin Mo, Nathan Mo, Richard Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究诊断多跳问答中的事实锚定差距,发现失败分为检索失败和提取失败,后者占近半缺陷且无法通过检索改进解决,需不同解决方案。

中文摘要 AI 辅助

多跳问答需要结合多个文档中的信息来回答复杂问题。这些系统已变得日益强大,然而当它们失败时,错误通常被归因于未能找到正确的文档。这种情况在单个推理步骤层面上是否成立,在很大程度上仍未得到检验。我们跨三个标准多跳问答基准对此进行了研究,并发现失败可分解为两种不同的模式:检索失败,即所需段落未被检索到;以及提取失败,即段落已被检索到但所需事实无法被提取——我们将这一现象称为事实锚定差距。提取失败占所有逐跳缺陷的近一半,并且对标准检索指标不可见。它们在我们测试的每一种检索干预下仍未得到解决,从而为仅靠检索的改进设定了一个上限。该差距的严重程度因基准和问题类型而异,但提取失败出现在我们测量的每个数据集上。我们的发现揭示,检索失败和提取失败是根本不同的瓶颈,需要不同的解决方案——而这一区别在当前评估实践中是缺失的。

英文摘要

Multi-hop question answering requires combining information from multiple documents to answer complex questions. These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not finding the right documents. Whether this holds at the level of individual reasoning steps remains largely unexamined. We investigate this across three standard multi-hop QA benchmarks and find that failures decompose into two distinct modes: retrieval failures, where the needed passage was not retrieved, and extraction failures, where the passage was retrieved but the needed fact could not be extracted - a phenomenon we term the fact-grounding gap. Extraction failures account for nearly half of all per-hop deficiencies and are invisible to standard retrieval metrics. They remain unresolved by every retrieval intervention we test, establishing a ceiling for retrieval-only improvements. The gap's severity varies across benchmarks and question types, but extraction failures appear on every dataset we measure. Our findings reveal that retrieval failures and extraction failures are fundamentally different bottlenecks requiring different solutions - a distinction absent from current evaluation practice.

发表机构

  • Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑