DocPC:基于代表性页面组合的文档级视觉检索
DocPC: Document-Level Visual Retrieval via Representative Page Composition
AI总结:
针对现有视觉文档检索以页面为中心的问题,提出DocPC框架,通过组合代表性页面实现文档级索引,在DocViRe基准上性能优于最强页面级基线且存储成本大幅降低。
AI中文摘要:
视觉文档检索通过使用视觉语言模型对页面截图进行编码,绕过了光学字符识别(OCR)流程,取得了进展。然而,现有方法仍以页面为中心,与需要完整文档检索的实际场景不匹配。简单的页面到文档聚合方法存在线性索引成本,且当相关性跨多个页面时检索性能会下降。我们提出DocPC,这是一种基于代表性页面组合的文档级视觉检索框架:选择代表性页面并将其组合成单个网格图像以进行文档级索引,这使得索引图像、向量和存储量减少了10.1倍,端到端索引时间减少了约7.7倍。为处理文档级普遍存在的多正样本监督,我们将多正样本对比学习与稀疏调度的列表式优化相结合。我们还引入了DocViRe,这是一个带有多正样本相关性标注的基准。DocPC-ColQwen在DocViRe上实现了44.09的NDCG@5,优于最强的页面级基准(38.91),同时减少了10.1倍的存储。代码和数据可在指定URL获取。
英文摘要:
Visual document retrieval has advanced by encoding page screenshots with vision-language models, bypassing OCR pipelines. However, existing methods remain page-centric, misaligned with real-world scenarios requiring complete document retrieval. A naive page-then-document aggregation suffers from linear indexing cost and degraded retrieval when relevance spans multiple pages. We propose DocPC, a document-level visual retrieval framework based on Representative Page Composition: selecting representative pages and composing them into a single grid image for document-level indexing, reducing indexed images, vectors, and storage by 10.1x and end-to-end indexing time by roughly 7.7x. To handle multi-positive supervision prevalent at the document level, we combine multi-positive contrastive learning with sparsely scheduled listwise optimization. We also introduce DocViRe, a benchmark with multi-positive relevance annotations. DocPC-ColQwen achieves NDCG@5 of 44.09 on DocViRe, outperforming the strongest page-level baseline at 38.91 while reducing storage by 10.1x. Code is available at https://anonymous.4open.science/r/DocPC-Document-Level-Visual-Retrieval-via-Representative-Page-Composition-1D52. Data is available at https://huggingface.co/datasets/anonymous-7219/docpc.