arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合成数据能将泰文OCR提升到何种程度?

How Far Can Synthetic Data Take Thai OCR?

Kunat Pipatanakul

arXiv 2609.03595首次发表:更新:

发表机构

Wayu Research; Paxa Labs(瓦尤研究所; 帕萨实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究合成OCR监督迁移到真实泰文文档的关键因素,构建Wayu-Paxa-OCR-Zero模型,仅用合成数据训练便大幅降低泰文OCR字符错误率,性能优于Typhoon OCR v1 7B。

AI 中文摘要

本研究探究合成OCR监督学习为何能迁移到真实泰文文档,并利用所得见解构建Wayu-Paxa-OCR-Zero,这是一款无需真实泰文文档OCR标签即可适配的泰文OCR模型。合成数据可提供大规模的精确标签,但“真实感”混淆了源域、页面上下文、排版、空间结构和字形变化等因素。我们通过可控文档重建流程对这些因素进行解耦,并在页面级和裁剪级训练下,针对印刷体和手写体泰文文档评估各变体的性能。非文本上下文的影响几乎无一致性,而字体多样性、二维结构和真实手写字形可提升迁移效果;此外,源域匹配依赖于训练粒度,在页面级训练下,域内重建的中位字符错误率为1.82%,接近真实印刷体监督的1.31%,但在裁剪级训练下,其性能逊于域外重建(15.59%对比5.52%)。基于这些发现,我们使用45723个合成页面将参数规模为0.9B的PaddleOCR-VL-1.6适配为Wayu-Paxa-OCR-Zero;相对于其基础检查点,该模型在印刷体页面上的中位字符错误率从6.64%降至1.24%,在手写体上从74.87%降至20.55%,且在全部五个评估集上均优于Typhoon OCR v1 7B,表明仅用合成数据训练即可具备竞争力。

英文摘要

We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction pipeline and evaluate each variant under page- and crop-level training on printed and handwritten Thai documents. Non-text context has little consistent effect, whereas typeface diversity, two-dimensional structure, and real handwriting glyphs improve transfer; moreover, source-domain matching depends on training granularity, with in-domain reconstruction approaching real printed supervision under page-level training (1.82% versus 1.31% median character error rate) but underperforming out-of-domain reconstruction under crop-level training (15.59% versus 5.52%). Guided by these findings, we adapt the 0.9B-parameter PaddleOCR-VL-1.6 into Wayu-Paxa-OCR-Zero using 45,723 synthetic pages: relative to its base checkpoint, it reduces median character error rate from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting and outperforms Typhoon OCR v1 7B on all five evaluation sets, showing that synthetic-only training can be competitive.

Comments20 pages, technical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑