MIDR:面向多模态文档检索的富集增强索引
MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval
- Bloomberg(彭博公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出无需训练的MIDR框架,将多模态推理移至索引时,在多领域文档检索任务中,其性能优于或匹敌现有模型,且索引内存和查询延迟更优。
AI中文摘要:
视觉丰富文档的检索存在表征问题:重要内容常存在于表格、图表、图形及布局关系中,而普通OCR会将其线性化,导致信息损坏或丢失。ColPali系列视觉检索器通过补丁级多向量索引和后期交互评分解决了该问题,使图像衍生的检索保留在查询时服务路径中。我们提出MIDR(Multimodal Indexing for Document Retrieval,面向文档检索的多模态索引),这是一种无需训练的富集增强索引框架,将多模态推理转移到索引时。在数据摄入阶段,多模态LLM将渲染后的页面转换为经验证的文本字段,这些字段通过BM25F索引,并可选择性与密集检索融合,实现基于多模态 grounding 证据的以文本为中心的服务。在ViDoRe V3数据集的五个英语领域中,MIDR Hybrid的平均nDCG为0.6219,相比BM25实现23.0%的相对提升,且与ColQwen2.5保持竞争力。在两个法语文档领域,富集技术缩小了英语查询与法语页面文本的差距,将BM25的nDCG从0.1532提升至0.5448,且性能优于ColQwen2.5。在全部七个领域中,MIDR在四个领域优于ColQwen2.5,同时索引内存约小9倍,查询延迟约低2倍。这些结果表明,索引时多模态推理是服务时视觉后期交互的一种极具吸引力的精度-部署替代方案。
英文摘要:
Retrieval over visually rich documents has a representation problem: important content often lives in tables, charts, figures, and layout relations that plain OCR linearizes, corrupts, or omits. ColPali-family visual retrievers address this with patch-level multi-vector indexes and late-interaction scoring, keeping image-derived retrieval on the query-time serving path. We introduce MIDR (Multimodal Indexing for Document Retrieval), a training-free framework for enrichment-augmented indexing that shifts multimodal reasoning to index time. During ingestion, a multimodal LLM converts rendered pages into verified textual fields that are indexed with BM25F and optionally fused with dense retrieval, enabling text-centric serving over multimodally grounded evidence. On ViDoRe V3, MIDR Hybrid achieves 0.6219 average nDCG across five English domains, a 23.0% relative gain over BM25, remaining competitive with ColQwen2.5. On two French-document domains, enrichment bridges English queries and French page text, lifting BM25 from 0.1532 to 0.5448 nDCG and outperforming ColQwen2.5. Across all seven domains, MIDR leads ColQwen2.5 on four while using approximately 9x smaller index memory and approximately 2x lower query latency. These results establish index-time multimodal reasoning as a compelling accuracy-deployment alternative to serving-time visual late interaction.