arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25140cs.CVcs.CL

RefLAM:一种用于历史阿拉伯文手稿的参考依据型行标注流水线

RefLAM: A Reference-Grounded Line Annotation Pipeline for Historical Arabic Manuscripts

发表机构高等计算机科学学院(ESI-SBA)
查看机构详情
  • Higher School of Computer Science (ESI-SBA)(高等计算机科学学院(ESI-SBA))

机构由 AI 辅助整理,请以论文原文为准。

Mohamed Guechaoui, Mohamed Diaa Zellagui, Souleyman Chaib, Sahraoui Dhelim

首次发表
浏览论文内容

中文总结 AI 辅助

RefLAM是用于历史阿拉伯文手稿的参考依据型行标注流水线,结合深度学习、MLLM与模糊对齐引擎,可高效生成标注数据,吞吐量较手动标注提升75倍,发布的AraMS-28k可用于HTR模型微调。

中文摘要 AI 辅助

现有的构建阿拉伯文手写文本识别(HTR)训练数据的行级方法,要么依赖无法扩展的完全手动标注,要么依赖尚未扩展到多脚本、双区域(正文加边注)手稿布局且无可靠正确性保证的自动OCR-参考对齐方法。本文提出RefLAM(手稿的参考依据型行标注),这是一种将手稿页面图像和干净转录文本转换为经验证的行级真值的流水线,同时不牺牲人工监督。RefLAM结合了深度学习页面分割模型、用于结构化OCR的多模态大语言模型(MLLM),以及与元音符号无关的模糊对齐引擎,该引擎将每条OCR行对应到参考文本的连续跨度,并给出[0,100]范围内的字符级置信度分数。完美分数在验证后可证明等价于归一化字符串的逐字符一致性(即置信度-100规则),在发布的语料库中未发现反例。因此,评审人员可信任完美分数,无需重新输入即可快速确认大多数行,使标注成为分级过程,注意力集中在不确定的对齐上。在7本经完整页面验证的书籍中,我们测量到其吞吐量比手动标注提高了75倍(3000行/小时 vs 40行/小时);对另外7本应用相同保证后,我们在一周内保留了16533条置信度-100的正文行,排除了分数低于100的行而非手动修正它们。通过RefLAM,我们发布了AraMS-28k:14本历史阿拉伯文手稿书籍、3043个页面,以及27971条正文行和629条边注行标注,带有边界框、布局标签和191个边注条目的插入锚点(占比30.4%)。我们还在AraMS-28k上微调了经Muharaf预训练的基线模型(包括HATFormer),并报告了字符错误率(CER)结果,证实其对下游HTR训练的实用价值。

英文摘要

Existing approaches to building line-level Arabic handwritten-text-recognition (HTR) training data either rely on fully manual annotation, which does not scale, or on automatic OCR-to-reference alignment methods not yet extended to multi-script, two-zone (main-plus-margin) manuscript layouts with a provable correctness guarantee. We present RefLAM (Reference-grounded Line Annotation for Manuscripts), a pipeline converting manuscript page images and clean transcriptions into validated, line-level ground truth without sacrificing human oversight. RefLAM couples a deep-learning page-segmentation model with a multimodal large language model (MLLM) for structured OCR and a diacritic-agnostic fuzzy alignment engine that grounds each OCR line in a contiguous span of the reference text, with a character-level confidence score in $[0,100]$. A perfect score is provably equivalent to character-for-character identity of the normalised strings (the Confidence-100 rule), verified with no counterexample across the released corpus. A reviewer can thus trust a perfect score, confirming most lines at a glance rather than retyping them, so annotation becomes triaged, with attention concentrated on uncertain alignments. Across 7 fully page-validated books we measured a 75$\times$ throughput gain over manual annotation (3,000 vs. 40 lines/hr); applying the same guarantee to 7 further books, we retained 16,533 confidence-100 main-text lines within one week, excluding sub-100 lines rather than manually correcting them. Using RefLAM, we release AraMS-28k: 14 historical Arabic manuscript books, 3,043 pages, and 27,971 main-text and 629 margin-line annotations with bounding boxes, layout labels, and insertion anchors for 191 margin entries (30.4%). We also finetune Muharaf-pretrained baselines (including HATFormer) on AraMS-28k and report CER results confirming its practical utility for downstream HTR training.

补充信息

↑