arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

波斯像素:一个用于波斯语的大规模合成OCR数据集

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

Pouria Mahdi, Haq Nawaz Malik

arXiv 2607.20385首次发表:更新:

AI 中文总结

研究针对波斯语OCR不太成熟的问题,介绍波斯像素数据集。该数据集用SynthOCR - Gen框架从波斯语料库生成超34.3万个图像文本对,经多种退化模型增强,为训练OCR架构提供资源,推动波斯语OCR发展。

AI 中文摘要

尽管有超过1.1亿人在多个国家说波斯语,但波斯语的光学字符识别(OCR)仍远不如拉丁文字语言成熟。这一差距源于两个基本挑战:波斯 - 阿拉伯文字系统的内在复杂性以及大规模、高质量注释数据集的有限可用性。本文介绍了波斯像素,一个专门设计用于应对这些挑战的综合合成OCR数据集。它包含超过34.3万个高保真图像文本对,跨越句子、段落和全页文档布局,由精心策划的七百万字波斯语语料库使用SynthOCR - Gen渲染框架生成。生成管道忠实地模拟了波斯文字的排版特征。为弥合合成到真实领域的差距,渲染图像进一步通过二十多种随机退化模型进行增强。波斯像素克服了注释波斯OCR数据长期稀缺的问题,为训练和微调现代OCR架构提供了可扩展且开放可用的资源,为波斯文档分析等研究奠定了坚实基础,同时表明程序化合成数据生成是推进低资源和排版复杂脚本OCR的实用、经济高效且可扩展的替代手动注释的方法。

英文摘要

Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the limited availability of large-scale, high-quality annotated datasets. Persian script exhibits obligatory cursive connectivity, context-dependent glyph shaping, extensive ligatures, diacritic placement, and stylistic variation across writing forms such as Naskh and Nastaliq, all of which significantly complicate text recognition. At the same time, the high cost and labor-intensive nature of manual annotation have created a persistent data bottleneck, limiting the development of robust OCR systems and slowing progress in Persian document digitization.In this paper, we introduce Persian Pixel, a comprehensive synthetic OCR dataset specifically designed to address these challenges. Comprising over 343,000 high-fidelity image text pairs, the dataset spans sentence, paragraph, and full-page document layouts generated from a carefully curated seven-million-word Persian corpus using the SynthOCR-Gen rendering framework. The generation pipeline faithfully models the typographic characteristics of Persian script, including contextual character joining, positional glyph variants, diacritic placement, and multiple representative Persian typefaces. To bridge the synthetic-to-real domain gap, the rendered images are further enriched with more than twenty-five stochastic degradation models that emulate realistic document acquisition artifacts, including ink bleed, paper aging, blur, illumination variation, scanner imperfections, compression artifacts, and multiple noise processes.By overcoming the long-standing scarcity of annotated Persian OCR data, Persian Pixel provides a scalable and openly available resource for training and fine-tuning modern OCR architectures, including transformer-based models such as TrOCR and Donut. The dataset establishes a strong foundation for research in Persian document analysis, historical manuscript digitization, and end-to-end document understanding, while demonstrating that programmatic synthetic data generation offers a practical, cost-effective, and scalable alternative to manual annotation for advancing OCR in low-resource and typographically complex scripts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑