发表机构
Université de Montréal; Concordia University(蒙特利尔大学; 康考迪亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AdaMerge通过检测合并余弦序列中的悬崖边界,实现无需每数据集调参的补丁压缩,在保持检索精度的同时显著降低存储和延迟,并优于现有方法。
AI 中文摘要
多向量视觉文档检索(VDR)模型(如ColPali和ColNomic)通过将每个文档表示为数百至数千个补丁级嵌入来获得强大的准确性,但代价是巨大的存储和延迟开销。现有的压缩方法要么修剪不重要的补丁,要么将相似补丁合并为簇;最近的最先进的合并方法Prune-then-Merge(PtM)在高压缩率下持续优于仅修剪的基线,但需要针对每个数据集通过网格搜索调整簇预算m。我们观察到,层次聚类产生的合并余弦序列呈现出陡峭的悬崖,将可合并的冗余与显著信号分开,并且该悬崖的位置在来自14个数据集的超过11,000个文档中集中在一个狭窄的带内。这表明合并边界可以针对每个文档检测,而不是针对每个数据集调整。基于这一观察,我们提出了AdaMerge,一种即插即用的压缩方法,它(i)通过合并余弦轨迹上的间隙分析检测每个文档自身的悬崖,以及(ii)构建注意力加权的簇质心以保留显著信号。在长文档基准ViDoRe-V2(4个数据集,两个骨干网络)上,AdaMerge在操作范围内显著优于调整后的PtM(p < 10^-4);在短文档基准ViDoRe-V1(10个数据集,两个骨干网络)上,所有合并方法已经接近无损,AdaMerge无需任何每数据集调整即可匹配调整后的PtM。AdaMerge每个文档仅增加约10毫秒,并暴露一个跨所有数据集和骨干网络共享的单一全局超参数。
英文摘要
Multi-vector visual document retrieval (VDR) models such as ColPali and ColNomic achieve strong accuracy by representing each document with hundreds to thousands of patch-level embeddings, at substantial storage and latency cost. Existing compression methods either prune unimportant patches or merge similar ones into clusters; the recent state-of-the-art merging method Prune-then-Merge (PtM) consistently outperforms pruning-only baselines at high compression, but requires a per-dataset cluster budget m to be tuned by grid search. We observe that the merge-cosine sequence produced by hierarchical clustering exhibits a sharp cliff separating mergeable redundancy from salient signal, and that the location of this cliff is concentrated in a narrow band across more than 11,000 documents from 14 datasets. This suggests the merge boundary can be detected per document rather than tuned per dataset. Building on this observation, we propose AdaMerge, a plug-and-play compression method that (i) detects each document's own cliff via gap analysis on the merge-cosine trajectory, and (ii) builds attention-weighted cluster centroids to preserve salient signal. On the long-document benchmark ViDoRe-V2 (4 datasets, two backbones), AdaMerge significantly outperforms tuned PtM across the operating range (p < 10^-4); on the short-document benchmark ViDoRe-V1 (10 datasets, two backbones), where all merging methods are already near-lossless, AdaMerge matches tuned PtM without any per-dataset tuning. AdaMerge adds only about 10 ms per document and exposes a single global hyperparameter shared across all datasets and backbones.
Comments5 pages, 3 figures. Accepted as a short paper at ACM CIKM 2026