arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越识别:结合候选选择分析与证据保留审核的紧凑多域阿拉伯手稿手写文本识别(HTR)

Beyond Recognition: Compact Multi-Domain Arabic Manuscript HTR with Candidate-Selection Analysis and Evidence-Preserving Review

Abdullah Ahmed Ali, Mohammed Thamer Abdulhadi, Ali Haider Safaa, Dhulfiqar Mahdi Wadi

arXiv 2608.19385首次发表:更新:

发表机构

University of Technology(理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Phoenix识别器与Athar审核工作流,通过多域适配技术提升阿拉伯手稿HTR性能,结果表明手稿HTR应作为可审计证据管理而非单纯文本替换进行评估。

AI 中文摘要

历史阿拉伯手稿转录并非单纯的识别问题,可用的学术系统必须应对书写者与版式的变化、保留不确定的读取结果、区分视觉证据与语言合理性,并记录研究者的最终决策。我们提出Phoenix,一个含499万参数的CNN-BiLSTM-CTC识别器,以及围绕其构建的证据感知审核工作流Athar。Phoenix通过文档感知重放、扩展至81个符号的编解码器,以及拒绝以不可接受的先前域性能损失换取新域性能提升的检查点的遗忘防护,适配档案、马格里布及历史手稿域。在与评估前冻结的先前检查点的指定保留集对比中,Phoenix在10594行Agapet文本上将字符错误率(CER)从22.12%降至17.86%,在11684行Omar文本上从17.72%降至11.84%,但在164行TariMa文本上从10.39%升至10.72%;在两个大型保留集上,字符加权CER从19.98%降至14.93%,相对误差降低25.3%。采用相同协议的单独开发诊断显示,Phoenix在4个可比域中均达到最低CER(未加权宏观CER为9.59%)。N-best诊断显示,束搜索解码与Oracle@25间存在2.15分的神谕差距,而神经文本重排序器、共识最小贝叶斯风险(MBR)、CTC后验质量估计及局部预CTC隐状态质量估计仅能弥补该差距的不到4%。因此,Athar保留视觉读取结果、呈现有限的替代选项、保守使用局部语言模型、检索具有独特、模糊或弃权(不执行)状态的来源平行文本,并导出可审计的TEI与PAGE-XML记录。研究结果支持将手稿HTR评估视为可审计的证据管理,而非单纯的文本替换。

英文摘要

Historical Arabic manuscript transcription is not only a recognition problem. A usable scholarly system must cope with shifting hands and layouts, preserve uncertain readings, distinguish visual evidence from linguistic plausibility, and record the researcher's final decision. We present Phoenix, a 4.99-million-parameter CNN-BiLSTM-CTC recognizer, and Athar, an evidence-aware review workflow built around it. Phoenix is adapted across archival, Maghrebi, and historical manuscript domains using document-aware replay, an expanded 81-symbol codec, and forgetting guards that reject checkpoints that improve a new domain at unacceptable cost to previous domains. In a pre-specified held-out comparison against the preceding checkpoint, frozen before evaluation and scored with greedy decoding and raw references, Phoenix reduced CER from 22.12% to 17.86% on 10,594 Agapet lines and from 17.72% to 11.84% on 11,684 Omar lines, while regressing from 10.39% to 10.72% on 164 TariMa lines. Across the two large held-out sets, character-weighted CER fell from 19.98% to 14.93%, a 25.3% relative error reduction. A separate same-protocol development diagnostic found the lowest CER for Phoenix on four of four comparable domains (9.59% unweighted macro CER). An N-best diagnostic revealed a 2.15-point oracle gap between beam decoding and Oracle@25, while neural text rerankers, consensus MBR, CTC-posterior quality estimation, and local pre-CTC hidden-state quality estimation recovered less than 4% of this gap. Athar therefore preserves the visual reading, exposes bounded alternatives, uses local language models conservatively, retrieves source parallels with unique, ambiguous, or abstain states, and exports auditable TEI and PAGE-XML records. The results support evaluating manuscript HTR as auditable evidence management rather than silent text replacement.

Comments13 pages, 4 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑