AI 中文总结
该研究针对向量数据库中检索枢纽问题,提出查询感知准入索引机制,发现其维护成本存在观测限制下限,通过召回率感知哨兵探测机制降低开销,在pgvector中实现且摄入开销仅0.33%。
AI 中文摘要
在服务于生产规模检索的向量数据库中,单篇插入文档可能会在查询工作负载中占据异常大的检索份额,成为检索枢纽,并主导整个主题返回的证据。一种新兴防御机制在数据摄入阶段通过准入检查来应对此问题:它维护一组哨兵查询,仅当文档针对这些查询的反向k近邻(reverse-kNN)数量保持在阈值τ以下时,才允许该文档被摄入。在工作负载漂移场景下,这组哨兵是必须在线维护的查询感知辅助索引,我们研究了这种维护对摄入路径造成的成本。我们发现了一个结构限制:覆盖并非冗余——当某个区域被覆盖后,监控器会停止提升哨兵,但谓词仅在τ个哨兵观测到枢纽时才会拒绝该枢纽,因此暴露度存在一个受观测限制的下限,无论减少更新或执行延迟都无法消除该下限。在基于880万向量的MS MARCO语料库上的真实HNSW、IVF-Flat和IVF-PQ索引上,该下限仅为最优情况:随着索引召回率下降,暴露度和 churn( churn指数据更新或状态变化的频率)会上升至该下限以上,且在召回率低于约0.5时,该门控机制会完全失去对枢纽的约束——在十亿级规模下使用的内存压缩IVF-PQ上情况最糟;而感知召回率的哨兵探测机制可将约束性恢复,且准入成本为固定的O(|S|d),仅为ANN插入操作的0.1%以下。我们在真实(COVID-19)工作负载漂移下验证了该规律,在PostgreSQL/pgvector中实现了该门控机制,其摄入开销仅为0.33%,并将该边界转化为一个配置规则,可按新兴区域调整哨兵预算大小。计数测试可约束检索时间得分归一化器(NNN、QB-Norm)无法约束的枢纽,且预注册的因果套件将缺失覆盖机制与跨两个嵌入族(BGE-1024、E5-768)的检索碎片化隔离开来。
英文摘要
In a vector database serving production-scale retrieval, a single inserted document can be retrieved for an anomalously large share of the query workload -- a retrieval hub -- and dominate the evidence returned for an entire topic. An emerging defense guards against this at ingest with an admission check: it maintains a set of sentinel queries and admits a document only if its reverse-kNN count against them stays below a threshold tau. Under workload drift this sentinel set is a query-aware auxiliary index that must be maintained online, and we study the cost that maintenance imposes on the ingest path. We identify a structural limit -- coverage is not redundancy: a monitor stops promoting sentinels once a region is covered, but the predicate rejects a hub only once tau sentinels witness it, so exposure has an observation-limited floor that no reduction in update or enforcement latency can close. On real HNSW, IVF-Flat, and IVF-PQ indexes over an 8.8M-vector MS MARCO corpus this floor is only a best case: as index recall falls, exposure and churn rise above it, and below recall ~0.5 the gate stops containing altogether -- worst on the memory-compressed IVF-PQ used at billion scale -- while a recall-aware witness probe restores containment at a fixed O(|S|d) admission cost, under 0.1% of the ANN insert. We validate the law under real (COVID-19) workload drift, implement the gate in PostgreSQL/pgvector at a 0.33% ingest tax, and turn the bound into a provisioning rule that sizes the sentinel budget per emerging region. A count test contains the hub where retrieval-time score normalizers (NNN, QB-Norm) do not, and a pre-registered causal suite isolates the missing-coverage mechanism from retrieval fragmentation across two embedding families (BGE-1024, E5-768).
Comments8 figures. Preprint; under submission