发表机构
IBM Research(IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对BM25的语义词汇鸿沟问题,本文提出CE-QE方法,利用交叉编码器的语义检索结果扩展BM25查询,在多个BEIR数据集上显著提升检索性能且不修改底层索引。
AI 中文摘要
词汇检索(BM25)可捕获精确的关键词匹配,并根据语料库范围的重要性对术语进行加权,但它存在语义词汇鸿沟的缺陷:当相关文档对答案的表述与查询不同时,BM25永远无法检索到该文档,下游的任何重排序或融合都无法恢复从未出现在候选集中的文档。本文提出交叉编码器查询扩展(Cross-Encoder Query Expansion,CE-QE),该方法会读取应用于顶级语义搜索结果的交叉编码器的每个标记相关性归因,选择交叉编码器视为决定性的术语,并将其附加到BM25查询中。与复用BM25自身(可能错误的)顶级结果的经典伪相关反馈不同,CE-QE从语义检索器的结果中生成扩展,避免了自我强化的查询漂移;与近期的生成式查询扩展(HyDE、Query2doc)不同,后者会提示大型语言模型从其参数知识中生成虚构文本,而CE-QE的每个扩展术语都直接从检索到的段落中逐字复制,因此不会引入语料库中不存在的词汇,且其仅增加的成本是对混合管道已用于重排序的交叉编码器进行归因提取。在七个BEIR数据集上,CE-QE在查询与答案词汇差异较大的场景中大幅提升了词汇召回率(例如,NQ的Recall@100从0.32提升至0.47),其分数融合变体(SESF)在Recall@100上比交叉编码器分数融合高出2.5%,在nDCG@10上分别比SPLADEv2和ColBERTv2高出5.3%和4.6%,同时完全不修改底层的BM25索引。
英文摘要
Lexical retrieval (BM25) captures exact keyword matches and weights terms by corpus-wide significance, but it is blind to the semantic vocabulary gap: when a relevant document phrases an answer differently from the query, BM25 never retrieves it, and no amount of downstream reranking or fusion can recover a document that was never in the candidate set. We present Cross-Encoder Query Expansion (CE-QE), which reads the per-token relevance attributions of a cross-encoder applied to top semantic search results, selects the terms the cross-encoder treats as decisive, and appends them to the BM25 query. Unlike classical pseudo-relevance feedback, which reuses BM25's own (possibly wrong) top results, CE-QE seeds expansion from the semantic retriever's results, avoiding self-reinforcing query drift. Unlike recent generative query expansion (HyDE, Query2doc), which prompts a large language model to hallucinate text from its parametric knowledge, every CE-QE expansion term is copied verbatim from a retrieved passage, so it cannot introduce vocabulary the corpus does not contain, and its only added cost is attribution extraction on a cross-encoder a hybrid pipeline already runs for reranking. On seven BEIR datasets, CE-QE improves lexical recall substantially where query and answer vocabulary diverge (e.g., NQ Recall@100 from 0.32 to 0.47), and its score-fusion variant (SESF) beats cross-encoder score fusion by 2.5% on Recall@100 and beats SPLADEv2 and ColBERTv2 by 5.3% and 4.6% on nDCG@10, while leaving the underlying BM25 index completely unmodified.