视觉语言模型(VLMs)何时助力阿拉伯手稿光学字符识别(OCR)?一项跨数据集研究
When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study
浏览论文内容
中文总结 AI 辅助
本文通过跨8类阿拉伯语文本数据集的实验,探究VLMs对阿拉伯手稿OCR的助力场景,提出OCR先验可恢复性原则,支持构建自适应OCR-VLM工作流。
中文摘要 AI 辅助
视觉语言模型(VLMs)正越来越多地被用于文档理解,但它们在阿拉伯语及伊斯兰手稿识别中的作用仍未得到充分探索。为解决这一研究空白,本文评估了传统OCR、通用VLMs、阿拉伯语专用VLMs以及经OCR条件化的VLM校正方法,覆盖8个阿拉伯语文本数据集,涵盖历史手稿、老旧印刷书籍、清晰印刷文本、多领域文档及手写文本。结果显示,不存在在所有设置下均占优的单一方法:在行级历史手稿上,VLMs性能接近Tesseract;在页级手稿图像上,VLMs表现更优;在若干场景中,经OCR条件化的校正器性能优于独立OCR和独立VLMs。核心发现为OCR先验可恢复性原则:当OCR输出在视觉和文本层面仍可恢复时,OCR条件化可提供锚点,供VLMs结合图像进行优化,这一设置可提升老旧印刷文本、清晰印刷文本、混合领域阿拉伯语文本及部分纳斯赫(Naskh)手稿的识别性能;但当先验存在脚本不匹配或系统性误导时,如马格里布(Maghribi)手稿及真实学生手写文本,该设置会降低性能。额外诊断显示,阿拉伯语VLM-OCR对变音符号、预处理、生成预算及重复循环敏感。这些发现支持构建自适应OCR-VLM工作流,该工作流可根据脚本、OCR先验可恢复性、长度诊断及失败模式指标对页面进行路由处理。
英文摘要
Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remains underexplored. To address such a gap in this paper, we evaluate traditional OCR, general-purpose VLMs, Arabic-specialized VLMs, and OCR-conditioned VLM correction across eight Arabic text datasets spanning historical manuscripts, aged printed books, clean print, multi-domain documents, and handwriting. The results show that no single approach dominates across setups. On line-level historical manuscripts, VLMs are close to Tesseract; on page-level manuscript images, they perform better; and in several settings, an OCR-conditioned corrector improves over both standalone OCR and standalone VLMs. The central finding is an OCR-prior recoverability principle: OCR conditioning helps when the OCR output remains visually and textually recoverable, providing anchors that the VLM can refine against the image. It improves recognition on aged print, clean print, mixed-domain Arabic, and some Naskh manuscripts, but degrades performance when the prior is script-mismatched or systematically misleading, as in Maghribi manuscripts and realistic student handwriting. Additional diagnostics show that Arabic VLM-OCR is sensitive to diacritics, preprocessing, generation budget, and repetition loops. These findings support an adaptive OCR-VLM workflow that routes pages according to script, OCR-prior recoverability, length diagnostics, and failure-mode indicators.
发表机构
- Qatar Computing Research Institute(卡塔尔计算研究所)
- Hamad Bin Khalifa University(哈马德·本·哈利法大学)
- Northwestern University in Qatar(卡塔尔西北大学)
- University of Doha for Science Technology(多哈科技大学)
机构由 AI 辅助整理,请以论文原文为准。