Better Together: 强RAG基线下的互补查询改写
Better Together: Complementary Query Rewriting Under a Strong RAG Baseline
浏览论文内容
中文总结 AI 辅助
本研究证明查询改写应作为强RAG基线的互补覆盖源,通过成本感知路由组合多种策略,在企业数据上提升HIT@10达12.5点,而非独立替代基线。
中文摘要 AI 辅助
改进检索增强生成(RAG)的一种流行方法是将用户的问题改写为多个变体,并使用所有这些变体进行检索。我们测试了在底层检索已经很强的情况下,这种方法是否真的有帮助。在一个固定的、有竞争力的流水线(BGE稠密检索、交叉编码器重排序和MMR多样化)下,我们将四种查询改写策略(S1-S4)与两个强LLM基线(HyDE、Query2Doc)在三个数据集(HotpotQA、AmbigNQ和512K文档的EnterpriseRAG-Bench)上进行了比较,使用三个随机种子并进行配对自助显著性检验。我们的主要结果是,单独改写最多与强基线相当,但组合方法能带来超额的收益,因为不同的策略在不同的问题上失败。四种方法(S1+S3+S4+HyDE)的事后并集在HIT@10上比基线提高了+12.5个百分点(51.70对39.22),五种方法的并集达到52.98(+13.8)。预算匹配的对照仅捕获了约40%的收益,证实互补性而非检索预算是主要驱动力。在HotpotQA上,并集增加了+1.6到+1.8个百分点(p<0.001),达到全方法预言机的饱和;在AmbigNQ上,同样的融合反而有害(低于最佳单方法-2.4,p<0.001),我们分析了何时以及为何如此。由于改写成本高昂,我们在模拟中评估了一个置信度门控路由器,仅在基线自身的top-1得分较低时才运行改写。它捕获了企业全合并收益的约一半(+4.3 HIT@10),同时仅在<40%的查询上支付改写成本,并在AmbigNQ上自动拒绝改写。下游答案质量评估证实,路由器在约40%的扩展成本下将F1提高了+1.92(p<0.01)。总之:将查询改写视为互补的覆盖来源,通过成本感知路由应用,而不是作为强基线的独立替代品。
英文摘要
A popular way to improve Retrieval-Augmented Generation (RAG) is to rewrite the user's question into several variants and search with all of them. We test whether this actually helps once the underlying search is already strong. Under one fixed, competitive pipeline (BGE dense retrieval, cross-encoder reranking, and MMR diversification), we compare four query-rewriting strategies (S1-S4) against two strong LLM baselines (HyDE, Query2Doc) on three datasets (HotpotQA, AmbigNQ, and the 512K-document EnterpriseRAG-Bench) over three seeds with paired-bootstrap significance tests. Our headline result is that rewriting alone is at best competitive with a strong baseline, but combining methods yields outsized gains because different strategies fail on different questions. A post-hoc union of four methods (S1+S3+S4+HyDE) improves HIT@10 over the baseline by +12.5 points on enterprise data (51.70 vs 39.22), and a five-method union reaches 52.98 (+13.8). Budget-matched controls capture only ~40% of this gain, confirming that complementarity, not retrieval budget, is the primary driver. On HotpotQA the union adds +1.6 to +1.8 points (p<0.001), saturating the all-method oracle; on AmbigNQ the same fusion hurts (-2.4 below the best solo, p<0.001), and we analyze when and why. Because rewriting is expensive, we evaluate in simulation a confidence-gated router that runs rewriting only when the baseline's own top-1 score is low. It captures about half of the enterprise full-merge gain (+4.3 HIT@10) while paying rewriting cost on <40% of queries, and automatically declines to rewrite on AmbigNQ. A downstream answer-quality evaluation confirms the router improves F1 by +1.92 (p<0.01) at roughly 40% of the expansion cost. In short: treat query rewriting as a complementary coverage source, applied through cost-aware routing, not as a standalone replacement for a strong baseline.
发表机构
- ServiceNow
机构由 AI 辅助整理,请以论文原文为准。