arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hi-Q:用于多跳问答的分层证据引导查询细化方法

Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering

Jueun Kim, Sungho Park, Wook-Shin Han

arXiv 2608.30468首次发表:更新:

发表机构

POSTECH(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多跳问答中问题与证据粒度不匹配的瓶颈,提出基于证据条件的分层查询细化框架Hi-Q,在全语料检索等设置下,其性能优于IRCoT、PropRAG等基线方法。

AI 中文摘要

多跳问答(QA)的核心瓶颈在于,问题表述的粒度通常与语料库证据可检索的粒度不匹配。现有方法通过在语料库上施加固定图结构、迭代重构查询或在语料库上执行生成的程序来解决这一问题,但这些策略未明确决定何时查询单元已得到证据支持、何时应细化查询。我们将该瓶颈形式化为可检索粒度发现问题,并提出Hi-Q——一种基于证据条件的分层查询细化框架。在每个查询节点,一个解析算子会测试检索到的证据是否支持当前查询单元;已解析节点终止,未解析节点则通过保留依赖关系的二元算子展开,并由语义覆盖验证器检查。因此,Hi-Q会生成一个查询树,其拓扑结构由语料库支持信号决定,而非固定分解模板或预构建图。我们在三个多跳QA基准上评估Hi-Q,主要采用全语料检索设置,该设置下依赖证据需在开放域干扰项中定位,而非小型标注池中。此设置中,Hi-Q在三个基准上平均达到52.3的精确匹配(EM)和64.0的F1值,在该平均值上比迭代检索基线IRCoT高出15.1 EM/18.2 F1,且无需构建全语料图,在MuSiQue-full上比基于图的基线PropRAG高出11.5 EM/12.0 F1。在过往研究使用的受限支持/干扰项设置中,Hi-Q同样取得最佳准确率,平均达到57.9 EM和69.3 F1,比PropRAG高出5.6 EM/3.9 F1,比IRCoT高出13.7 EM/15.8 F1。项目页面可访问该https URL。

英文摘要

A central bottleneck in multi-hop Question Answering (QA) is that the granularity at which a question is expressed often differs from the granularity at which corpus evidence is retrievable. Existing methods address this mismatch by imposing fixed graph structures over the corpus, by iteratively reformulating the query, or by executing a generated program over it, but these strategies do not explicitly decide when a query unit is already supported by evidence and when it should be refined. We formulate this bottleneck as retrievable granularity discovery and introduce Hi-Q, an evidence-conditioned framework for hierarchical query refinement. At each query node, a resolution operator tests whether retrieved evidence supports the current query unit; resolved nodes terminate, while unresolved nodes are expanded by a dependency-preserving binary operator and checked by a semantic coverage verifier. Hi-Q therefore grows a query tree whose topology is determined by corpus support signals rather than by a fixed decomposition template or a pre-built graph. We evaluate Hi-Q on three multi-hop QA benchmarks, primarily under full-corpus retrieval, where dependent evidence must be located among open-domain distractors rather than within a small annotated pool. In this setting Hi-Q reaches 52.3 EM and 64.0 F1 averaged over the three benchmarks, ahead of the iterative retrieval baseline IRCoT by 15.1 EM / 18.2 F1 on that same average, and ahead of the graph-based RAG baseline PropRAG by 11.5 EM / 12.0 F1 on MuSiQue-full, without corpus-wide graph construction. In the restricted supporting/distractor setting used by prior work, Hi-Q likewise attains the best accuracy, with 57.9 EM and 69.3 F1 on average, ahead of PropRAG by 5.6 EM / 3.9 F1 and IRCoT by 13.7 EM / 15.8 F1. The project page is available at https://hi-q-project.github.io/.

Comments28 pages, 9 figures. Project page: https://hi-q-project.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑