arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11553cs.IR

EVIE:用于视觉文档检索的证据向量感知嵌入

EVIE: Evidence-Vector-Informed Embeddings for Visual Document Retrieval

  • Tencent IMA Product Center(腾讯智能多媒体产品中心)
  • Tencent Youtu Lab(腾讯优图实验室)

机构由 AI 辅助整理,请以论文原文为准。

Zifei Wang, Wei Wen, Qiang Ji, Qian-Wen Zhang, Ruizhi Qiao, Xing Sun

AI总结:

提出EVIE系列视觉文档检索器,通过三项创新解决现有VDR方法的局限性,在138项任务上验证,提升了VDR的准确性与存储效率的权衡。

AI中文摘要:

准确且可扩展的视觉文档检索(VDR)既需要细粒度的页面理解,又需要高效的索引,但现有方法难以同时满足这两点。基于OCR的文本检索会增加预处理延迟,且可能丢失理解复杂页面所需的视觉和结构线索;单向量视觉-语言模型绕过了OCR,但将整个页面压缩为单个向量会限制查询与文档匹配的细粒度;采用MaxSim的多向量检索器能提供更细粒度的交互,但需要大型索引,且在准确性提升方面仍有空间。我们认为,克服这些局限性需要在表示学习和索引构建过程中保留与查询相关的页面证据。为此,我们提出了EVIE(Evidence-Vector-Informed Embeddings,证据向量感知嵌入)系列原生视觉文档检索器,整合了三项关键创新:(1)证据判断数据治理,利用多模态判别器识别含答案的正样本并过滤不可靠的负样本;(2)带对称列表式蒸馏和前缀式套娃表示学习(Prefix-MRL)的双向师生学习,使单个学生检查点可支持六个嵌套嵌入维度,无需重新编码;(3)层次聚合索引压缩(HAC),通过空间正则化对页面token进行聚类,并存储语义质心以实现单阶段MaxSim检索。在ViDoRe V1、V2、V3和JinaVDR的138项任务上开展的大量实验验证了EVIE的有效性:EVIE-8B在V3上达到66.75的nDCG@10,超出最佳外部基线1.43个点,四个套件的平均值为79.51;带HAC的EVIE-4.5B在每百万页仅3.81 GiB的情况下,仍保持59.58的nDCG@10,将向量负载降低了128倍。这些结果共同改善了视觉文档检索的准确性-存储权衡。

英文摘要:

Accurate and scalable visual document retrieval (VDR) requires both fine-grained page understanding and efficient indexing, yet existing approaches struggle to achieve both. OCR-based text retrieval adds preprocessing latency and can lose visual and structural cues needed to understand complex pages. Single-vector vision-language models bypass OCR, but compressing an entire page into one vector limits the granularity of query--document matching. Multi-vector retrievers with MaxSim provide finer interactions, yet demand large indexes and still leave room for accuracy improvements. We argue that overcoming these limitations requires preserving query-relevant page evidence throughout representation learning and index construction. To this end, we introduce \textbf{\textit{EVIE}} (Evidence-Vector-Informed Embeddings), a family of native visual document retrievers integrating three key innovations: (1) Evidence-judged data governance, which uses a multimodal judge to identify answer-bearing positives and filter unreliable negatives. (2) Bidirectional teacher--student learning with symmetric listwise distillation and prefix-based Matryoshka representation learning (Prefix-MRL), enabling one student checkpoint to serve six nested embedding dimensions without re-encoding. (3) Hierarchical agglomerative index compression (HAC), which clusters page tokens with spatial regularization and stores semantic centroids for single-stage MaxSim retrieval. Extensive experiments across 138 tasks from ViDoRe V1, V2, V3, and JinaVDR validate EVIE. EVIE-8B achieves 66.75 nDCG@10 on V3, exceeding the best external baseline by 1.43 points, with a four-suite average of 79.51. EVIE-4.5B with HAC retains 59.58 nDCG@10 at only 3.81 GiB per million pages, reducing vector payload by $128\times$. Together, these results improve the accuracy--storage trade-off for visual document retrieval.

补充信息

↑