arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20064cs.CV

免费的午餐?将 PP-OCRv6 适配用于历史文本识别

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

Benjamin Kiessling

首次发表
浏览论文内容

中文总结 AI 辅助

本文评估了轻量级识别器 PP-OCRv6 在历史文本识别中的表现,发现其通过异构预训练和微调可优于传统 CRNN 及大型视觉语言模型,兼具准确性与实用性。

中文摘要 AI 辅助

尽管大型视觉语言模型在报告中的得分令人印象深刻,但由于其计算成本高、依赖大规模预训练以及存在幻觉问题,它们在历史自动文本识别中的实际应用仍然有限。因此,历史 ATR 在很大程度上仍依赖于紧凑的 CRNN 行识别器,这些识别器具有视觉基础,并且可以在适度的数据上进行训练。轻量级无循环识别器有望在具备 CRNN 实际优势的同时达到更大模型的准确性,但尚未在历史书写中得到全面评估。我们将 PP-OCRv6(一种近期推出的、无强语言建模的紧凑文本识别器)适配用于历史行识别,并在多语言拉丁字母和阿拉伯字母材料上,将其与传统 CRNN 在广义预训练、领域特定训练、语料库级微调和手稿特定少样本适配等场景下进行比较。虽然 PP-OCRv6 在从头训练时并未持续优于基线,但异构预训练产生了显著更好的泛化能力。与基于 Qwen3.5 的 Medusa 识别器的比较进一步表明,微调后的 PP-OCRv6 可以胜过针对历史拉丁字母 HTR 定制的大型 VLM。

英文摘要

Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CRNN line recognizers, which are visually grounded and trainable on modest data. Lightweight recurrence-free recognizers promise the accuracy of larger models with the practical advantages of CRNNs, yet have not been comprehensively evaluated on historical writing. We adapt PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compare it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material. While PP-OCRv6 does not consistently outperform the baseline when trained from scratch, heterogeneous pretraining produces markedly better generalization. Comparisons with the Qwen3.5-based Medusa recognizer further show that fine-tuned PP-OCRv6 can outperform a large VLM tailored towards historical Latin-script HTR.

发表机构

  • Inria Paris(法国国家信息与自动化研究所巴黎中心)

机构由 AI 辅助整理,请以论文原文为准。

↑