发表机构
Independent Researcher(独立研究员)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对马耳他语OCR资源匮乏的问题,提出合成训练管道和5流Tesseract LV-ROVER集成,在422段落基准上将字符错误率从0.0234降至0.00700(降幅70%),其中集成识别贡献44%的改进。
AI 中文摘要
马耳他语拥有不错的文本语料库和预训练语言模型,但与少数拥有大型OCR基准的语言不同,它只有一个已知的真实标注PDF语料库用于OCR训练,共57页,远低于段落级训练所需:属于OCR领域的低资源语言。由于没有真实的语料库进行大规模训练,我们构建了一个合成训练管道和一个5流Tesseract LV-ROVER集成,并在一个422段落基准上报告了结果,与微调Tesseract基线相比,字符错误率(CER)为0.0234。仅集成识别就将CER降低了44%,达到0.01317;一个五阶段后处理链将整个管道的CER降至0.00700,降低了70%。该链的大部分是印刷体规范化,但其中一个阶段是恢复误读的变音符号而非对齐标点,因此我们将其报告为识别增益,而不是将整个链归为一个标签。我们将44%的数值视为识别器所学内容的可迁移估计,而70%的数值则特定于该基准的标签约定。
英文摘要
Maltese has substantial text corpora and pretrained language models, but paragraph-scale OCR training data remains scarce; NOMOCRAT provides 57 verified annotated pages. LV-ROVER-MLT combines synthetic fine-tuning of Tesseract~5 with five complementary recognition streams and lexicon-gated word-level arbitration adapted to Maltese diacritics and hyphenation. In the DocEng~2026 Maltese OCR competition, the system placed first with held-out CER 0.0074; the next-ranked submission scored 0.0161 and NOMOCRAT scored 0.0163. The same approach produced a significant improvement over stock Tesseract on Luxembourgish, while the Hungarian result was inconclusive. A 36,803-pair Maltese OCR corpus constructed from EUR-Lex and Wikipedia provides an additional paragraph-level resource. Code, model weights, and corpus data are public.
Comments10 pages, including 8 page and references plus appendices. Working paper. Updated with DocEng 2026 Maltese OCR competition results; LV-ROVER-MLT placed first with held-out CER 0.0074