arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结合合成数据与真实数据的低资源历史OCR:满文案例研究

Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study

Yan Hon Michael Chung, Hanlin Wang

arXiv 2609.11495首次发表:更新:

发表机构

The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过结合合成与真实数据训练OCR模型,将满文历史文档识别准确率从87.4%提升至96%以上,并利用投票和词典规则进一步达到98.27%。

AI 中文摘要

满文如今已极度濒危,曾是清朝(1636-1912)的主要语言之一,其大量档案记录正日益数字化,但仍难以大规模检索和分析。先前的研究表明,仅使用合成满文单词图像训练的视觉-语言模型(VLM)在真实清代手稿和印刷品上可达到87.4%的单词准确率,但仍存在显著的合成到真实的差距。本研究探讨了在低资源OCR中应如何结合合成与真实历史训练数据。使用60,000张合成和20,306张真实历史单词图像,我们在四种训练模式下评估了三个预训练VLM和一个紧凑型卷积循环神经网络(CRNN):仅合成、仅真实、合成与真实联合训练、以及合成到真实的顺序训练,并遵循常见的检查点选择和档案评估协议。引入真实训练图像将领先配置的单词准确率提升至95.09%至96.28%之间,而所有仅合成配置均未超过87.92%。合成数据补充显著提升了所有三个VLM的性能,而其对CRNN的边际效果则对训练目标敏感。在测试的实际流程下,联合训练和顺序训练在档案准确率上大致相当。一旦获得真实图像,紧凑型CRNN也能达到领先性能范围,表明模型规模本身并不决定识别准确率。最后,强识别器之间的互补错误使得投票无需额外训练即可将准确率提升至98.27%,而一部十八世纪的满文词典为裁决分歧提供了原则性规则。

英文摘要

Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work showed that vision-language models (VLMs) trained only on synthetic Manchu word images can reach 87.4% word accuracy on real Qing manuscripts and prints, leaving a substantial synthetic-to-real gap. This study examines how synthetic and real historical training data should be combined for low-resource OCR. Using 60,000 synthetic and 20,306 real historical word images, we evaluate three pretrained VLMs and a compact convolutional recurrent neural network (CRNN) under four regimes: synthetic-only, real-only, joint synthetic-real, and sequential synthetic-to-real training, following a common checkpoint-selection and archival evaluation protocol. Introducing real training images raises the leading configurations to between 95.09% and 96.28% word accuracy, while no synthetic-only configuration exceeds 87.92%. Synthetic supplementation substantially improves all three VLMs, whereas its marginal effect for the CRNN is sensitive to the training objective. Joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines. A compact CRNN also reaches the leading performance range once real images are available, showing that model scale alone does not determine recognition accuracy. Finally, complementary errors among strong recognizers allow voting to raise accuracy to 98.27% without additional training, while an eighteenth-century Manchu dictionary provides a principled rule for adjudicating disagreements.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑