覆盖性很重要:用于压缩多向量视觉文档检索器的MarginMerge
Coverage Matters: MarginMerge for Compressing Multi-Vector Visual Document Retrievers
浏览论文内容
中文总结 AI 辅助
本文针对多向量视觉文档检索器压缩问题,提出MarginMerge方法,在6个数据集上实现高检索精度保留,大幅减少存储向量,且可跨数据集迁移。
中文摘要 AI 辅助
ColPali、ColQwen等多向量视觉文档检索器通过存储细粒度的图像块嵌入实现了强大的检索能力,但这会产生庞大的索引并带来高昂的晚期交互评分成本。本文认为,有效的压缩应保留与查询相关的覆盖性,即跨查询可能成为最强MaxSim匹配的不同文档区域,而非按显著性独立选择图像块,这一观点也解释了为何密集渲染页面比自然图像更易压缩。本文提出MarginMerge,一种用于冻结多向量检索器的压缩方法:它选择覆盖感知锚点,对文档图像块进行聚类,并使用轻量级共享网络为每个聚类合成一个代表。压缩在索引期间执行一次,检索时保持标准MaxSim接口。在ColQwen2.5和ColPali的六个数据集上,MarginMerge在5%和10%向量保留率下实现了最高的查询无关平均匹配值;与使用相同主干的未压缩索引相比,它保留了97%至99%的平均nDCG@5,同时减少了90%至95%的存储文档向量;在5%保留率下,它在所有六个ColQwen2.5数据集上相对于几何合并的排名翻转平均减少了约41%;该模型无需重新训练即可迁移到未见数据集和不同保留率。
英文摘要
Multi-vector visual document retrievers such as ColPali and ColQwen achieve strong retrieval by storing fine-grained patch embeddings, but this produces large indexes and costly late-interaction scoring. We argue that effective compression should preserve query-relevant coverage, meaning the diverse document regions that may become the strongest MaxSim match across queries, rather than selecting patches independently by salience. This view also explains why dense rendered pages are easier to compress than natural images. We introduce MarginMerge, a compression method for frozen multi-vector retrievers. It selects coverage-aware anchors, clusters document patches, and uses a lightweight shared network to synthesize one representative per cluster. Compression is performed once during indexing, while retrieval keeps the standard MaxSim interface. Across six datasets on both ColQwen2.5 and ColPali, MarginMerge achieves the highest matched query-agnostic average at 5% and 10% vector retention. Compared with the uncompressed index using the same backbone, it preserves between 97% and 99% of average nDCG@5 while reducing stored document vectors by between 90% and 95%. At 5% retention, it also reduces ranking flips relative to geometric merging on all six ColQwen2.5 datasets by approximately 41% on average. The same model transfers to unseen datasets and retention ratios without retraining.