arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迭代微调对复杂历史梵文手稿转录准确率的影响

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis

arXiv 2608.18696首次发表:更新:

发表机构

Centre for Inter-disciplinary Artificial Intelligence (CAI); FLAME University(跨学科人工智能中心(CAI); 火焰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对复杂历史梵文手稿的OCR挑战,提出可在布局与外观层级迭代微调的传统OCR流水线,构建含精细标注的数据集,验证了迭代微调的性能提升并完成多模态大模型基准测试。

AI 中文摘要

将手写历史手稿中的文本数字化,可使其更易获取、便于保存,并能让历史学者以新方式开展研究。然而,历史手稿往往因特定时期的书写风格、页面纹理、相机噪声及其他干扰因素,呈现出复杂多样的布局与非标准外观,这使得对其进行光学字符识别(OCR)颇具挑战。为应对这一挑战,我们提出一种局部传统OCR流水线,该流水线可在布局层级与外观层级上针对目标手稿进行迭代微调。通过适配目标手稿的分布,所提出的传统OCR流水线能对后续页面做出更准确的预测,从而逐步减少人工标注工作量——人工标注需具备历史领域专业知识,成本高昂且耗时。我们利用该流水线对三份复杂的历史梵文手稿进行文本数字化,并引入了一个包含精细布局层级标注及标准PAGE-XML格式Unicode标注的数据集。我们展示了所提出的传统OCR流水线经迭代微调后在量化性能上的提升,同时在该引入的数据集上对领先的多模态大语言模型的性能进行了基准测试。代码与数据集可在此httpsURL获取。

英文摘要

Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historical scholars to study them in new ways. Historical manuscripts, however, often exhibit complex heterogeneous layouts and non-standard appearance due to period-specific writing styles, page textures, camera noise, and other nuisance factors, making them difficult to perform OCR on. To tackle this challenge, we introduce a local traditional OCR pipeline, which can be iteratively fine-tuned on the target manuscript at the layout-level and the appearance-level. By adapting to the target manuscript distribution, the proposed Traditional OCR pipeline makes better predictions on subsequent pages, causing iterative reduction in human annotation effort, which is expensive and time-consuming as it requires historical domain expertise. Using this pipeline, we digitize text from three complex historical Sanskrit manuscripts and introduce a dataset with granular layout-level annotations, along with Unicode annotations in the standard PAGE-XML format. We demonstrate quantitative gains due to iterative fine-tuning of the proposed traditional OCR pipeline, and also benchmark the performance of leading Multi-Modal Large Language Models on the introduced Dataset. Code and dataset are available at: https://github.com/flame-cai/gnn-synthetic-layout-historical/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑