arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16309cs.IRcs.AI

跨稀疏检索机制的静态剪枝:可迁移的、失效的及仍有效的部分

Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps

Zirui Song, Yuye Zhu, Yang Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过跨三种检索引擎的1140组实验,明确索引侧静态剪枝可迁移,查询剪枝失效,静态剪枝可与动态剪枝互补,还给出了可迁移的剪枝停止准则。

中文摘要 AI 辅助

静态剪枝被广泛用于加速稀疏神经检索,但现有研究均在单一自定义流程内验证其结论,导致哪些发现可迁移至具有不同索引组织和动态剪枝机制的现代引擎尚不明确。我们开展首个跨引擎剪枝可迁移性研究,在三种引擎(受控C++流程:倒排索引;BMP:块最大剪枝;SEISMIC:聚类倒排索引)上评估静态剪枝策略,在两个基准(MS MARCO、Natural Questions)上使用两种覆盖相反查询密度区间的编码器(SPLADE:平均44个查询词;V3-GTE:平均7个查询词),共1140种实验配置,另在TREC DL 2019/2020上开展深度判断验证。我们发现索引侧剪枝(文档及倒排列表)可迁移:因稀疏检索受内存限制,其在所有引擎上均持续降低延迟(1.2-6.6倍)和索引大小(18%-82%),该结论得到缓存未命中、TLB及IPC分析支持。相反,查询剪枝已被现代引擎内化:其在穷尽流程上可实现4-11倍加速,但已被BMP的β和SEISMIC的query_cut所覆盖。静态剪枝与动态剪枝互补:在BMP上,结合文档与查询缩减可实现2.5倍加速,且NDCG@10与精确基线的差值在0.003以内。最终,在所有三种引擎上,NDCG@10趋于饱和时Recall@10仍处于约85%-95%区间,提供了可迁移的停止准则:从业者可将剪枝推进至该拐点,而不会出现明显的排序性能下降。综上,这些发现明确了可迁移的部分(索引侧剪枝)、失效的部分(查询剪枝)及仍有效的部分(动态剪枝之上的静态剪枝)。

英文摘要

Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single custom pipeline, leaving it unclear which findings transfer to modern engines with different index organizations and dynamic pruning mechanisms. We present the first cross-engine pruning portability study, evaluating static pruning strategies across three engines - a controlled C++ pipeline (exhaustive inverted index), BMP (block-max pruning), and SEISMIC (clustered inverted indexes) - on two benchmarks (MS MARCO, Natural Questions) with two encoders spanning opposite query-density regimes (SPLADE: 44 avg. query terms; V3-GTE: 7 avg. query terms), totaling 1,140 experimental configurations, with an additional deep-judgment validation on TREC DL 2019/2020. We find that index-side pruning (document and posting-list) is portable: it consistently reduces latency (1.2-6.6$\times$) and index size (18-82%) across all engines because sparse retrieval is memory-bound - a conclusion we support with cache-miss, TLB, and IPC profiling. In contrast, query pruning is already internalized by modern engines: it yields 4-11$\times$ speedup on the exhaustive pipeline but is subsumed by BMP's $β$ and SEISMIC's query_cut. Static pruning complements dynamic pruning: on BMP, combining document and query reduction yields 2.5$\times$ speedup with NDCG@10 within 0.003 of the exact baseline. Finally, NDCG@10 saturates while Recall@10 is still in the ${\sim}$85-95% range across all three engines, providing a portable stopping criterion: practitioners can push pruning to this knee without visible ranking degradation. Together, these findings answer what transfers (index-side pruning), what breaks (query pruning), and what still helps (static atop dynamic pruning).

发表机构

  • Amazon Web Services(亚马逊网络服务)

机构由 AI 辅助整理,请以论文原文为准。

↑