发表机构
Kyung Hee University; Korea University(庆熙大学; 高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有视觉文档检索模型以英语为中心的局限,提出韩文视觉文档单向量检索模型KoVRE,采用708729对双语查询-页面对训练,2B参数模型在基准测试中表现优于同类模型。
AI 中文摘要
视觉文档检索(Visual Document Retrieval,VDR)直接将文本查询与文档图像匹配,保留了文本提取过程中可能丢失的视觉和结构信息。然而,现有的VDR模型和训练资源仍以英语为中心,且许多高性能系统依赖大规模主干网络或存储密集型多向量表示。为解决这些局限,我们提出KoVRE(Korean Visual Document Retrieval Embedding),这是一种针对韩文视觉文档的单向量检索器,同时配套完整的训练方案。我们使用708729对韩文及英文查询-页面对训练该模型,采用正样本感知的难负样本挖掘技术,并对训练数据构成、难负样本处理以及基于重排序器的知识蒸馏开展控制分析。在韩文视觉文档检索基准测试中,我们的2B参数模型较基础主干模型有显著提升,其性能超过了8B参数的单向量对应模型以及一个强大的多向量基线模型。这些结果表明,针对性的双语监督和我们精心设计的训练策略,能够在不扩大主干网络或采用多向量表示的情况下,生成适用于不同文档领域的高效韩文VDR模型。
英文摘要
Visual Document Retrieval (VDR) directly matches text queries against document images, preserving visual and structural information that may be lost during text extraction. However, existing VDR models and training resources remain predominantly English-centric, while many high-performing systems rely on massive backbones or storage-intensive multi-vector representations. To address these limitations, we introduce KoVRE: Korean Visual Document Retrieval Embedding, a single-vector retriever for Korean visual documents, alongside a comprehensive training recipe. We train the model on 708,729 Korean and English query-page pairs using positive-aware hard-negative mining and conduct controlled analyses of training-data composition, hard-negative treatment, and reranker-based knowledge distillation. Across Korean visual document retrieval benchmarks, our 2B model substantially improves over the base backbone model, outperforming both its 8B single-vector counterpart and a strong multi-vector baseline. These results demonstrate that targeted bilingual supervision and our carefully designed training strategies can produce a highly effective Korean VDR model across diverse document domains, without requiring a scaled-up backbone or multi-vector representations.