CrossQ:用于晚期交互检索的任务对齐跨令牌条件量化
CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction Retrieval
AI总结:
CrossQ是一种任务对齐跨令牌条件量化方法,通过利用索引时的轻量文档上下文优化令牌编码,在压缩晚期交互检索器索引的同时,提升了检索性能,优化了内存-性能权衡。
AI中文摘要:
ColBERT等晚期交互检索器虽能达到较高质量,但存在多向量索引规模大的问题。标准压缩方法最小化令牌重构误差,而检索性能关键取决于保留稀疏“获胜”令牌的分数。本文提出CrossQ,通过在索引时计算轻量文档上下文(不存储),对令牌编码进行条件化,自适应提升文档内的有效令牌保真度。CrossQ采用与检索排名对齐的目标函数训练,以保留候选分数分布并保护难负样本边界。在每个令牌2比特(2 B/token)的设置下,CrossQ比最强的严格匹配内存占用的量化基线提升MRR@10达0.010,比最强的候选匹配系统参考提升MRR@10达0.012;在9个数据集组成的BEIR子集上,每个令牌4比特(4 B/token)时,CrossQ比最强的候选匹配系统参考提升平均nDCG@10达0.009。每个令牌4比特时,CrossQ实现了64倍的原始令牌存储缩减,若计入元数据约为61倍,按保守的填充/对齐计算约为58倍;每个令牌8比特(8 B/token)时,经轻量微调的CrossQ保留了全精度ColBERT约98%的MRR@10,优化了内存受限场景下晚期交互检索的内存-性能权衡。
英文摘要:
Late-interaction retrievers like ColBERT achieve high quality but suffer from large multi-vector indices. Standard compression minimizes token reconstruction error, while ranking depends critically on preserving scores of sparse "winner" tokens. We introduce CrossQ, which adaptively improves effective token fidelity within documents by conditioning token codes on lightweight document context computed at indexing time (but not stored). CrossQ is trained with ranking-aligned objectives that preserve candidate score distributions and protect hard-negative margins. At 2 B/token, CrossQ improves MRR@10 by +0.010 over the strongest strictly footprint-matched quantization baseline and by +0.012 over the strongest candidate-matched system reference. On a nine-dataset BEIR subset, CrossQ improves average nDCG@10 by +0.009 at 4 B/token over the strongest candidate-matched system reference. At 4 B/token, CrossQ achieves 64x raw token-storage reduction, approximately 61x including metadata and approximately 58x under conservative padding/alignment accounting. At 8 B/token, CrossQ with light fine-tuning retains approximately 98% of full-precision ColBERT MRR@10, improving the footprint-quality tradeoff for memory-constrained late-interaction retrieval.