arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EviQE:基于LLM的查询扩展的证据选择

EviQE: Evidence Selection for LLM-Based Query Expansion

Hai Son Le, Amin Bigdeli, Shirin Seyedsalehi, Morteza Zihayat, Ebrahim Bagheri

arXiv 2609.14875首次发表:更新:

发表机构

Toronto Metropolitan University; University of Waterloo; University of Toronto(多伦多都会大学; 滑铁卢大学; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EviQE通过聚合多个改写器的检索文档并选择紧凑证据集进行查询扩展,在TREC DL和BEIR基准上,基于相关性的证据选择(LLM-Score)显著优于直接改写等方法。

AI 中文摘要

基于LLM的查询扩展越来越多地根据从目标语料库中检索到的文档来调整改写,但大多数工作关注如何生成扩展,而非模型应读取哪些文档。我们提出EviQE,它聚合多个改写器检索到的文档,选择一个紧凑的证据集,并将其用于一次基于证据的扩展步骤。这将证据选择与生成分离,并将改写器视为互补的检索视角。在三个TREC DL和五个BEIR基准上,改写器经常检索到不同的相关文档,因此池化候选比任何单一来源提供更高的相关文档覆盖率。最强的增益来自基于相关性的证据选择:LLM-Score始终优于直接改写、冷启动扩展和单源种子扩展。一旦选择了强条件证据,额外的检索-生成轮次带来的收益很小,甚至可能降低效果。

英文摘要

LLM-based query expansion increasingly conditions reformulation on documents retrieved from the target corpus, yet most work focuses on how to generate expansions rather than which documents the model should read. We propose EviQE, which aggregates documents retrieved by multiple reformulators, selects a compact evidence set, and uses it for one grounded expansion step. This separates evidence selection from generation and treats reformulators as complementary retrieval perspectives. Across three TREC DL and five BEIR benchmarks, reformulators frequently retrieve distinct relevant documents, so pooled candidates provide higher relevant-document coverage than any individual source. The strongest gains come from relevance-based evidence selection: LLM-Score consistently outperforms direct reformulation, cold-start expansion, and single-source seeded expansion. Additional retrieval-generation rounds provide little benefit once strong conditioning evidence has been selected and can reduce effectiveness.

DOI:10.1145/3799682.3839965

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑