CoRECT:大规模评估嵌入压缩技术的框架
CoRECT: A Framework for Evaluating Embedding Compression Techniques at Scale
AI总结:
CoRECT框架通过大规模评估和新数据集,展示非学习压缩在减少索引大小的同时保持性能,为压缩方法的选择提供一致比较依据。
AI中文摘要:
密集检索系统在各种基准测试中已被证明是有效的,但需要大量内存来存储大型搜索索引。最近的嵌入压缩进展表明,通过最小的排名质量损失可以大幅减少索引大小。然而,现有研究往往忽视了语料复杂性的作用——这是一个关键因素,因为最新研究显示,语料规模和文档长度对密集检索性能有显著影响。在本文中,我们引入了CoRECT(压缩技术的受控检索评估框架),这是一个用于大规模评估嵌入压缩方法的框架,支持由新整理的数据集集合。为了展示其用途,我们基准测试了八种代表性压缩方法。值得注意的是,我们展示了非学习压缩在多达1亿条文档上也能实现显著的索引大小减少,且性能损失统计上不显著。然而,选择最佳压缩方法仍然具有挑战性,因为性能在不同模型间变化较大。这种变异性突显了CoRECT的必要性,以实现一致的比较和有指导的压缩方法选择。所有代码、数据和结果均可在GitHub和HuggingFace上获得。
英文摘要:
Dense retrieval systems have proven to be effective across various benchmarks, but require substantial memory to store large search indices. Recent advances in embedding compression show that index sizes can be greatly reduced with minimal loss in ranking quality. However, existing studies often overlook the role of corpus complexity -- a critical factor, as recent work shows that both corpus size and document length strongly affect dense retrieval performance. In this paper, we introduce CoRECT (Controlled Retrieval Evaluation of Compression Techniques), a framework for large-scale evaluation of embedding compression methods, supported by a newly curated dataset collection. To demonstrate its utility, we benchmark eight representative types of compression methods. Notably, we show that non-learned compression achieves substantial index size reduction, even on up to 100M passages, with statistically insignificant performance loss. However, selecting the optimal compression method remains challenging, as performance varies across models. Such variability highlights the necessity of CoRECT to enable consistent comparison and informed selection of compression methods. All code, data, and results are available on GitHub and HuggingFace.