arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22316cs.LGcs.CV

现代手写体热身是否有助于历史阿拉伯文OCR?针对Muharaf和KHATT的可复现、计算匹配评估

Does a Modern-Handwriting Warm-Up Help Historical Arabic OCR? A Reproducible, Compute-Matched Evaluation on Muharaf and KHATT

Sumaih Almarshad, Maram Alamri, Dona Aloraini, Fares Altuwaim, AlJawharh AlOtaibi, Reem Alyabis, Rayah Aldawsari

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过可复现、计算匹配的实验评估现代手写体热身对历史阿拉伯文OCR的影响,发现其并非普遍有效,发布了可独立验证的SaudiHeritage-OCR包。

中文摘要 AI 辅助

现代阿拉伯文手写体的中间阶段对历史阿拉伯文手写文本识别(HTR)是有帮助还是有害,通常仅通过一次实现和一次比较来判定,这对任何一方的主张而言依据过于薄弱。我们通过运行四次名义上相同的 ablation 实验来测试稳定性,在开发过程中自然变化基础检查点、编码器冻结策略、轮次预算、精度和学习率调度,同时保持归一化、评分器和区间估计固定。每次运行都比较在现代手写体(KHATT)上进行中间训练,然后在历史手稿(Muharaf)上微调,与直接在Muharaf上微调的效果。四次运行中,估计的效果在-17.64到+14.52个字符错误率(CER)点之间波动,甚至出现符号反转。两个极端值恰好对应存在可识别混淆因素的两次运行:一次学习率低5倍,另一次是来源未公开的检查点;两次干净运行的结果为-0.25和+0.94,即无效果。单一实现的紧密区间对下一次实现毫无意义。随后我们运行计算匹配实验,在三个随机种子上使用相同预算:KHATT热身比匹配的同域对照差+2.42个CER点(95%区间[+0.60, +4.25]);该差距中特定于手写体领域的部分仅约0.6个点,为此配置下的小负效应,并非普遍结果。我们发布了SaudiHeritage-OCR包,包含归一化器、区间评分器、经验证的KHATT解码器、实验清单、VLM基线和版本对齐协议,以便结果可独立验证。Al-Mahd铭文行被严格保留,未作为基准提供。

英文摘要

Whether an intermediate stage of modern Arabic handwriting helps or hurts historical Arabic HTR is usually decided from one implementation and one comparison, too thin a basis for a claim either way. We test stability by running the same nominal ablation four times, letting the base checkpoint, encoder-freezing strategy, epoch budget, precision, and learning-rate schedule vary as they naturally did during development, while holding the normalization, scorer, and interval estimation fixed. Each run compares intermediate training on modern handwriting (KHATT) then fine-tuning on historical manuscripts (Muharaf) against fine-tuning on Muharaf directly. Across the four runs the estimated effect swings from -17.64 to +14.52 CER points and reverses sign. The two extremes are exactly the two runs with an identifiable confound (a fivefold lower learning rate in one; a checkpoint of undisclosed provenance in the other); the two clean runs land at -0.25 and +0.94, i.e. no effect. A tight interval from one implementation says nothing about the next. We then run a compute-matched experiment with identical budgets over three seeds: KHATT warm-up is +2.42 CER points worse than a matched same-domain control (95% interval [+0.60, +4.25]); the part of that gap specific to the handwriting domain is only about 0.6 points a small negative effect under this configuration, not a universal result. We release a SaudiHeritage-OCR package with the normalizer, interval scorer, a verified KHATT decoder, experimental manifests, VLM baselines, and an edition-alignment protocol, so the result can be checked independently. The Al-Mahd inscription line is held strictly out and is not offered as a benchmark.

发表机构

  • Dal Research Team(达尔研究团队)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑