发表机构
Portland State University; EvenUp(波特兰州立大学; EvenUp)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RAG系统固定top-k检索的弊端,提出基于预检索查询聚类的自适应检索深度框架,离线估计查询难度并聚类推荐检索深度,在线常数时间分配,在真实流量中F1提升36%以上且token减少14%。
AI 中文摘要
RAG系统通常检索固定数量的文档(top-k)来为生成提供依据,但这种静态方法很脆弱:简单查询会遭受过度检索(增加噪声和成本),而复杂查询则检索不足,导致召回失败,进而级联为错误答案。受“必须检索多少文档才能可靠地回答任意查询”这一问题的驱动,我们提出了一种实用、通用的查询自适应检索深度框架。离线阶段,我们通过测量默认检索器下的NDCG来估计每个查询的检索难度,并从NDCG-k曲线中推导出查询特定的“饱和”点k*。由于在线计算这些信号代价高昂,我们在嵌入空间中对大量查询进行聚类,并使用均值加方差规则为每个聚类总结一个推荐检索深度,以目标高覆盖率(例如,约95%)。在运行时,系统将传入查询分配到一个聚类,并在常数时间内选择相应的top-k。与依赖检索文档聚类的检索后置信度方法相比,我们的方法是预检索且以查询为中心的,使其在异构、类似案例的语料库中具有鲁棒性,并适用于法律、医疗、金融和企业搜索等领域。最后,该框架已在全流量查询中测试,在低复杂度聚类上将F1提高了超过36%,同时将token使用量减少了14%,且没有准确性损失。
英文摘要
RAG systems commonly retrieve a fixed number of documents (top-k) to ground generation, but this static approach is brittle: simple queries suffer over-retrieval (adding noise and cost) while complex queries are under-retrieved, causing recall failures that cascade into incorrect answers. Motivated by the question of how many documents must be retrieved to answer an arbitrary query reliably, we propose a practical, general framework for query-adaptive retrieval depth. Offline, we estimate per-query retrieval difficulty by measuring NDCG under the default retriever and deriving a query-specific saturation point k* from the NDCG-k curve. Because computing these signals online is expensive, we cluster a large set of queries in embedding space and summarize each cluster with a recommended retrieval depth that targets high coverage (e.g., ~95%) using a mean-plus-variance rule. At runtime, the system assigns an incoming query to a cluster and selects the corresponding top-k in constant time. Compared with post-retrieval confidence methods that rely on clustering retrieved documents, our approach is pre-retrieval and query-centric, making it robust in heterogeneous, case-like corpora and applicable across domains such as legal, healthcare, finance, and enterprise search. Finally, this framework has been tested in full-traffic queries that improved $F_1$ by over 36% while reducing token usage by 14% on low-complexity clusters without accuracy loss.
CommentsAccepted to the Applied Research Track of CIKM 2026