AI 中文总结
该研究针对RAG系统固定Top-L截断融合的问题,提出EAHR方法,以完整列表加权RRF的Top-K为目标,通过PVS和PBM实现自适应精确检索,可复现完整列表Top-20并显著降低延迟。
AI 中文摘要
现代检索增强生成(RAG)系统常融合密集检索器与稀疏检索器的固定Top-L结果,将后续贡献视为零,该截断值同时决定排序与执行成本。但截断融合通常不等同于完整列表融合:即使观测候选包含完整列表Top-K的所有项,未读取的跨列表排序仍可能改变Top-K的成员资格或顺序。由于通道排序随查询和语料库更新而变化,从历史查询中选择的深度无法可靠迁移。我们提出精确自适应混合检索(EAHR),将完整列表加权RRF定义的有序Top-K作为检索目标,把通道深度视为特定请求的执行状态。逐向量标量量化(PVS)和倒排块最大值(PBM)生成可恢复的精确密集与稀疏排序,融合操作限制未读取贡献,仅在能改变Top-K时请求进一步排序,所有成功请求均匹配无预设Top-L的完整列表融合,否则安全执行至列表耗尽。在5个测试集合和5个时间语料快照中,完整列表加权RRF保持竞争力,而从历史查询选的深度无法可靠迁移;EAHR在全部150个查询-快照组合中复现了完整列表有序Top-20。在暖缓存、交错、顺序均衡协议下,穷尽批处理与EAHR的配对几何平均延迟比在TREC-DL 2019上为23.35,在TREC-DL 2020上为30.28;反相关排序会耗尽两个列表,部分困难查询用EAHR时更慢。EAHR不保证每个请求都提速,它在适配当前排序深度的同时确保结果精确。
英文摘要
Modern retrieval-augmented generation (RAG) systems often fuse fixed Top-$L$ results from dense and sparse retrievers, treating later contributions as zero. The cutoff therefore determines both the ranking and its execution cost. Yet truncated fusion is not generally equivalent to complete-list fusion: unread cross-list ranks can change Top-$K$ membership or order even when the observed candidates contain every item in the complete-list Top-$K$. Because channel rankings vary across queries and corpus updates, a depth selected from historical queries may not transfer reliably. We propose Exact Adaptive Hybrid Retrieval (EAHR), which fixes the ordered Top-$K$ defined by complete-list weighted RRF as the retrieval target and treats channel depth as request-specific execution state. Per-Vector Scalar Quantization (PVS) and Posting Block-Max (PBM) produce resumable exact dense and sparse rankings. Fusion bounds unread contributions and requests further ranks only while they can change the Top-$K$. Every successful request therefore matches complete-list fusion without a preset Top-$L$; otherwise, execution continues safely to list exhaustion. Across five test collections and five temporal corpus snapshots, complete-list weighted RRF remained competitive, whereas fixed depths selected from historical queries did not transfer reliably. EAHR reproduced the complete-list ordered Top-20 in all 150 query-snapshot combinations. Under a warm-cache, interleaved, order-balanced protocol, the paired geometric-mean latency ratios of exhaustive batch execution to EAHR were 23.35 on TREC-DL 2019 and 30.28 on TREC-DL 2020. Anti-correlated rankings exhausted both lists, and some difficult queries were slower with EAHR. EAHR does not guarantee a speedup for every request; it fixes the exact result while adapting execution depth to the current rankings.
Comments10 pages, 4 figures, 1 table