arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KoViDoRe:韩文视觉文档检索

KoViDoRe: Korean Visual Document Retrieval

Yongbin Choi, Yongwoo Song, Mujeen Sung

arXiv 2608.20840首次发表:更新:

发表机构

Kyung Hee University(庆熙大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有韩文视觉文档检索基准的不足,本文推出KoViDoRe基准及配套训练数据集Ko-VDR Train Public,评估发现现有模型难以处理该任务,为相关模型开发提供支撑。

AI 中文摘要

多模态检索的最新进展提升了从PDF和报告等视觉丰富文档中检索信息的能力,但现有基准大多以英语为中心,对结构复杂的韩文视觉文档覆盖有限;此外,多数现有韩文资源主要评估单页检索,无法捕捉需要跨多页聚合证据的现实场景。为解决这些差距,我们推出KoViDoRe——一个韩文视觉文档检索基准,该数据集由具有多样布局(包括表格、图表和多列结构)的公开韩文文档构建;我们开发了一个多阶段数据整理流程,包含结构化文档解析、基于摘要和上下文两种策略的合成查询生成,以及经人工验证的相关性映射。利用KoViDoRe,我们评估了多种多模态检索模型,发现当前模型难以有效处理韩文视觉文档检索,尤其是在涉及结构化内容和多样查询类型的场景中;受这一发现驱动,我们进一步整理了大规模训练数据集Ko-VDR Train Public,以支持针对韩文视觉文档的检索模型开发。KoViDoRe和Ko-VDR Train Public共同为韩文视觉文档检索提供了统一的基准和训练资源。

英文摘要

Recent advances in multimodal retrieval have improved the ability to retrieve information from visually rich documents such as PDFs and reports. However, existing benchmarks remain largely centered on English and provide limited coverage of Korean visual documents with complex structures. Furthermore, most existing Korean resources primarily evaluate single-page retrieval, failing to capture realistic scenarios that require evidence aggregation across multiple pages. To address these gaps, we introduce KoViDoRe, a benchmark for Korean visual document retrieval. The dataset is constructed from publicly available Korean documents with diverse layouts, including tables, figures, and multi-column structures. We develop a multi-stage data curation pipeline consisting of structured document parsing, synthetic query generation using both summary-based and context-based strategies, and relevance mapping with human verification. Using KoViDoRe, we evaluate a wide range of multimodal retrieval models and observe that current models struggle to effectively handle Korean visual document retrieval, particularly in settings involving structured content and diverse query types. Motivated by this finding, we further curate a large-scale training dataset, Ko-VDR Train Public, to support the development of retrieval models tailored to Korean visual documents. Together, KoViDoRe and Ko-VDR Train Public provide a unified benchmark and training resource for Korean visual document retrieval.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑