arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28248cs.CVcs.CL

Synth-JDoc:用于OCR的、包含多样布局与嵌入图像的日文文档图像数据集合成

Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images

Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建了含多样布局与嵌入图像的日文文档合成OCR数据集,可有效提升大型视觉语言模型读取竖排日文文本的性能,相关资源已公开。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)读取文档图像内文本的能力至关重要,这能支撑文档视觉问答等各类应用。为提升LVLMs的文本读取能力,高质量的OCR数据集必不可少,这对兼具竖排与横排文本的日文文档尤为关键。当前LVLMs对竖排日文文本的性能远低于横排,亟需专门的OCR数据集缩小差距。然而,手动构建OCR数据集成本高、难扩展;利用OCR模型从现有文档图像提取文本构建数据集,又会引入文本识别错误、需获取文档图像等问题。为解决这些问题,我们直接从文本合成文档图像构建OCR数据集:借助HTML和CSS生成包含竖排、横排的多栏文档,还在布局中嵌入文本到图像模型生成的图像以保证视觉真实性,同时添加噪声与退化滤波器提升合成文档图像的模型鲁棒性。实验中,我们将在本合成数据集上微调的模型,与在过往工作合成数据集、高性能文本到图像模型生成数据集上微调的基线模型对比,结果显示本合成数据集是提升LVLMs读取竖排日文文本性能的最有效方法,数据集与代码已公开。

英文摘要

The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. To enhance the text-reading capabilities of LVLMs, high-quality OCR datasets are essential. This need is particularly critical for Japanese documents, which often feature vertically written text alongside horizontally written text. Current LVLMs demonstrate considerably lower performance on vertically written Japanese text than on horizontally written text, necessitating specialized OCR datasets to bridge this gap. However, manually constructing OCR datasets is expensive and difficult to scale. Alternatively, constructing datasets by extracting text from existing document images using OCR models introduces challenges, such as text recognition errors and the prerequisite of sourcing document images. To address these issues, we construct an OCR dataset by synthesizing document images directly from text. Leveraging HTML and CSS, we generate multi-column documents that incorporate both vertical and horizontal writing styles. Furthermore, to ensure the visual realism of the documents, we embed images generated by text-to-image models within the layout. Additionally, to foster model robustness, we apply noise and degradation filters to the synthesized document images. In our experiments, we compared the performance of models fine-tuned on our synthetic dataset against baselines fine-tuned on synthetic datasets from prior work and those generated by a high-performance text-to-image model. Evaluation results demonstrate that our synthetic dataset is the most effective approach for improving LVLM performance on reading vertically written Japanese text. Our dataset and code are publicly available (https://github.com/llm-jp/synth-jdoc).

发表机构

  • Waseda University(早稻田大学)
  • NII(信息学研究所)
  • NII LLMC(信息学研究所语言媒体研究中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑