arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05441cs.DBcs.DCcs.ETcs.IR

分解式湖仓中的过滤向量搜索:表格式剪枝与逐文件ANN的组合

Filtered Vector Search in a Disaggregated Lakehouse: Composing Table-Format Pruning with Per-File ANN

Rakesh Jain, Thomas Griffin, Syed Zawad

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出在基于Apache Iceberg和Parquet的湖仓中,通过组合现有文件剪枝与逐文件IVF索引实现过滤向量搜索,在特定条件下可使ANN搜索速度提升数十倍。

中文摘要 AI 辅助

近似最近邻(ANN)搜索越来越多地与结构化数据协同运行——例如“找到租户为'acme'且语言为'en'的10个最近文档”——但相似度计算与过滤通常是分开实现的:一个是专门的向量索引,另一个是独立的过滤步骤。本文探讨当两者都存在于开放湖仓表(对象存储上基于Parquet的Apache Iceberg)中时会发生什么,该引擎已拥有成熟的文件剪枝栈(分区剪枝、区域映射、位图索引)。我们将IVF索引嵌入每个Parquet文件的页脚,通过组合表现有的文件剪枝与逐文件ANN来实现过滤向量查询的加速,而非采用新的过滤算法:规划器先通过谓词剪枝数据文件,再仅对剩余文件运行IVF。该索引以分布式、非破坏性方式构建——仅通过元数据层面的Iceberg替换,其他引擎仍可读取——且采用 rendezvous-hash 逐文件缓存,避免对象存储读取延迟抵消算法优势。收益完全来自文件剪枝:在1150万×768的表上,当recall@10≥0.90时,预热的IVF搜索比暴力搜索快约32倍,选择性谓词在ANN运行前已从444个数据文件中剪枝355个;在500万真实IBM Granite嵌入上,当过滤条件来自连接操作时,剪枝五分之四的区域分区后,在相同top-k下,比查询时连接快近两个数量级(94倍:14.7秒→157毫秒),前提是缩减操作已实现为区域分区布局。我们明确了该组合策略的适用条件:过滤列需具备文件级局部性,且剩余谓词仅可安全推入对可证明为纯(分区)列的搜索,而非仅排序的列;同时报告了将ANN集成到湖仓引擎时遇到的失败模式。

英文摘要

Approximate nearest-neighbor (ANN) search increasingly runs alongside structured data - "find the 10 nearest documents where tenant='acme' AND lang='en'" - yet similarity and filtering are usually bolted together: a specialized vector index for one, a separate filter step for the other. We ask what happens when both live inside an open lakehouse table (Apache Iceberg over Parquet on object storage), where the engine already owns a mature file-pruning stack (partition pruning, zone-maps, a bitmap index). We embed an IVF index in place in each Parquet file's footer and make filtered vector queries fast not with a new filtering algorithm but by composing the table's existing file pruning with per-file ANN: the planner prunes data files by the predicate first, then runs IVF only over the survivors. The index is built distributed and non-destructively - a metadata-only Iceberg replace that every other engine still reads - and a rendezvous-hashed per-file cache keeps object-store read latency from swamping the algorithmic win. The payoff comes entirely from file pruning. On an 11.5M x 768 table, warm IVF search is ~32x faster than brute force at recall@10 >= 0.90, a selective predicate having pruned 355 of 444 data files before ANN runs; on 5M real IBM Granite embeddings, a filter arriving across a join prunes four of five region partitions and runs nearly two orders of magnitude (~94x: 14.7 s -> 157 ms) faster than the query-time join at identical top-k, once the reduction is materialized into a region-partitioned layout. We characterize when the composition pays off - it requires file-level locality on the filter column, and the residual predicate is only safe to push into the search over a provably pure (partitioned) column, not a merely sorted one - and report the failure modes we hit bolting ANN onto a lakehouse engine.

↑