arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分层BM25:十亿文档规模的词汇搜索

Hierarchical BM25: Lexical Search at Billion-Document Scale

Umesh Deshpande, Swaminathan Sundararaman

arXiv 2608.00229首次发表:更新:

发表机构

IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对十亿文档规模下平面BM25索引内存占用大、查询延迟高的问题,提出分层BM25方法,通过粗粒度索引选择文档组,实现低延迟词汇检索,提升吞吐量并保持高召回率。

AI 中文摘要

针对十亿级文档的平面BM25索引约占400GB内存,若将其置于内存中,所需DRAM与语料库规模成正比;若从磁盘读取,每次查询耗时4-12秒,因此在交互式延迟预算内,该规模下的精确top-k词汇检索并不现实。分层BM25放弃精确排序,以换取内存和延迟的固定边界:常驻粗粒度索引会选择查询需访问的约1000个主题、大小均衡的文档组,采用两类信号,一是每个查询词在组内的总频率,二是针对跨组分布过疏、频率总和无法反映的信息词,需判断其中是否有多个词共同出现在同一文档中;随后对选中的组进行穷尽搜索,并基于约100KB的全局统计数据打分,所有返回的分数与平面索引的分数完全一致,近似仅体现在选择阶段。常驻内存占用约4.4GB,与语料库规模无关;针对十亿文档的16词查询耗时约300毫秒,吞吐量为平面多线程索引的4.7至5.6倍,预热后的缓存每秒可支持约32次查询,而平面索引每秒不足3次;在50万文档配置下,访问5%-10%的簇可恢复穷尽结果分数的0.83-0.92;十亿级召回及与文档重排序BlockMax-WAND的直接对比仍是待解决的问题。

英文摘要

A flat BM25 index over one billion documents occupies about 400 GB. Holding it in memory requires DRAM proportional to corpus size. Serving it from disk takes 4-12 seconds per query. Exact top-k lexical retrieval at this scale is therefore impractical within an interactive latency budget. Hierarchical BM25 gives up exact ranking in exchange for fixed bounds on memory and latency. A resident coarse index selects which of ~1K topical, size-balanced document groups a query visits, using two signals: the total frequency of each query term within a group, and, for informative terms spread too thinly across groups for frequency totals to reflect, whether several of them appear together in one document. Selected groups are then searched exhaustively and scored against ~100 KB of global statistics. Every returned score therefore equals the flat index's score, and the approximation is confined to selection alone. The resident footprint is ~4.4 GB, independent of corpus size. Sixteen-term queries over one billion documents return in ~300 ms (4.7x to 5.6x the throughput of a flat multi-threaded index), and a warmed cache sustains ~32 queries per second versus under 3 for flat indexing. At a 500K-document configuration, visiting 5-10% of clusters recovers 0.83-0.92 of the exhaustive result score. Billion-scale recall and a direct comparison against document-reordered BlockMax-WAND remain open.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑