生成式后期交互嵌入用于视觉文档检索
Generative Late-Interaction Embeddings For Visual Document Retrieval
浏览论文内容
中文总结 AI 辅助
针对视觉文档检索中后期交互嵌入存储开销大的问题,提出生成式后期交互嵌入(GLIE),利用向量几何特性从少量向量重建完整嵌入,在低存储下保持高检索精度。
中文摘要 AI 辅助
后期交互检索是视觉文档搜索的最先进技术,但其准确性以存储为代价。现有的压缩方法保留每页约1000个向量中的子集或局部平均值。然而,在激进的存储预算下,这些方法性能急剧下降,而替代方案需要重新训练编码器。通过跨三个编码器研究这种退化,我们发现两个一致的性质:向量精确地位于单位球面上,并集中在固有维度为五到六的流形附近。这一几何结构带来两点见解。第一,标准的k-means质心落在球体内部,导致MaxSim分数系统性低估。将它们归一化到球面是一种免费修正,相比原始质心可提升高达+0.093的nDCG@5。第二,由于页面流形自由度很少,完整向量集可以从少量向量重新生成。为此,我们引入了生成式后期交互嵌入(GLIE):每页k<<N个向量,从归一化质心学习,既作为轻量索引,又作为重新生成页面完整嵌入集的基础。在查询时,搜索仅在这些k个向量上进行,解码器仅将顶部候选扩展回所有N个向量以进行精确重新评分。在ViDoRe v1上每页四个向量时,GLIE保留了未压缩系统近80%的nDCG@5,而最佳先前后处理方法仅为70%。这些结果使用了415K参数的网络,在仅一千个训练页面上不到三分钟GPU时间内完成拟合。在匹配的训练预算下,微调编码器甚至达不到GLIE的无训练阶段,而完整系统在每个预算下都优于它。这些模式在第二个编码器和ViDoRe v2上保持一致。通过按需重建证据而非采样,GLIE为存储高效检索开辟了新维度,解码器作为其主要设计面。
英文摘要
Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means centroids fall inside the sphere, causing systematic underestimation of MaxSim scores. Normalizing them to the surface is a free correction worth up to +0.093 nDCG@5 over raw centroids. Second, because the page manifold has few degrees of freedom, the full set of vectors can be regenerated from only a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): k << N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set. At query time, search runs exclusively on these k vectors, and a decoder expands only the top candidates back to all N vectors for exact rescoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in under three GPU-minutes on just a thousand training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns hold across a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface.
发表机构
- King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
- Department of Computer Science, Edge Hill University(埃奇希尔大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。