发表机构
Amazon; Tel-Aviv University(亚马逊公司; 特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对向量搜索中固定参数导致的召回率方差及计算成本问题,提出QASP策略,通过前期回归预测召回曲线推导搜索策略,在99%召回率下减少80%数据访问,性能优异且可跨场景泛化。
AI 中文摘要
向量搜索的一项核心挑战是在最小化计算成本的同时实现始终较高的召回率。固定搜索参数会导致不同查询间的性能出现显著差异,而传统的平均召回率评估掩盖了这些逐查询的差异。我们提出QASP(Query-Adaptive robust vector Search Policy,查询自适应鲁棒向量搜索策略),该方法通过单次前期监督回归预测每个查询的完整召回率进展曲线,进而从中推导得出适用于任何召回目标的搜索策略;这避免了搜索过程中的迭代模型调用,也无需为每个单独的召回目标设置独立的预测器。QASP利用尺度不变特征和预搜索推理来预测归一化召回值,因此可跨召回目标、索引配置和数据集进行泛化。其细粒度的进展预测还支持轻量级的反应式补充机制,该机制可基于预测值与观测值的偏差调整搜索深度,且无需额外推理。我们证明,QASP所需的训练样本数量有限,与数据集规模和维度无关;其损失与任何固定策略的不可约下界之间的差值会逐渐消失;且相较于固定探测策略,它的数据访问量节省量随内在维度呈指数增长。实验表明,QASP实现了显著更低的召回率方差和与目标的偏差,更高的查询满意度,且无需重新训练即可扩展至大型数据和分层索引,在达到99%召回率的同时减少了80%的数据访问量。
英文摘要
A fundamental challenge of vector search is achieving consistently high recall while minimizing computational costs. Fixed search parameters cause significant performance variance across queries, and conventional evaluation on average recall masks these per-query disparities. We introduce QASP (Query-Adaptive robust vector Search Policy), which predicts the complete recall progression curve per query via a single upfront supervised regression, from which a search policy is derived for any recall target; this avoids iterative model invocations during search or separate predictors per target. By predicting normalized recall values with scale-invariant features and pre-search inference, QASP generalizes across recall targets, index configurations, and datasets. Its fine-grained progress predictions further enable a lightweight reactive complement that adjusts search depth based on predicted-versus-observed deviations without additional inference. We prove that QASP requires a finite training sample independent of dataset size and dimensionality, that its loss exceeds the irreducible lower bound of any fixed policy by a vanishing margin, and that its data access savings over fixed probing grow exponentially in intrinsic dimensionality. Experimentally, QASP achieves significantly lower recall variance and deviation from target, higher query satisfaction rate, and scales to large data and hierarchical indices without retraining, achieving 99% recall with 80% less data access.
Comments12 pages, 6 figures, 6 tables, preprint