发表机构
East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出OmniHandwritingOCR手写OCR诊断基准,评估13个多模态模型,发现其在复杂公式上表现差且存在幻觉,为诊断模型失效模式提供测试平台。
AI 中文摘要
多模态大语言模型(MLLMs)正越来越多地被用作文档和知识处理流程中的OCR系统,但它们准确读取真实手写文本的能力仍未得到充分探索。现有的OCR基准主要聚焦于印刷文本或干净的单行输入,对现实手写OCR场景的覆盖有限,例如多语言手写、书写者错误以及结构复杂的数学表达式。我们推出OmniHandwritingOCR,这是一个用于评估MLLMs和OCR系统在手写OCR任务中的诊断基准。它涵盖手写文本识别和手写数学表达式识别,分为六个子任务和十二个子集,总计77570张带标签图像,来自公开数据集和新收集的学生手写内容。一个关键组成部分是难度分层的多行公式语料库,旨在测试模型在结构复杂性不断增加时的鲁棒性。我们采用统一协议,使用五个互补指标评估了13个开源和闭源系统。结果表明,当前系统远未达到准确转录的水平:在复杂多行公式上的性能急剧下降,模型排名因语言和公式设置而异,且多个生成模型会产生看似合理但视觉上无依据的修正。OmniHandwritingOCR为诊断多模态模型在手写OCR场景中存在的语言、内容、结构和视觉 grounding(视觉接地)失效模式提供了具有挑战性的测试平台。
英文摘要
Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting remains underexplored. Existing OCR benchmarks focus largely on printed text or clean single-line inputs, leaving limited coverage of realistic handwritten OCR scenarios such as multilingual handwriting, writer errors, and structurally complex mathematical expressions. We introduce OmniHandwritingOCR, a diagnostic benchmark for evaluating MLLMs and OCR systems on handwritten OCR. It covers handwritten text recognition and handwritten mathematical expression recognition across six subtasks and twelve subsets, totaling 77.57K labeled images from public datasets and newly collected student writings. A key component is a difficulty-stratified multi-line formula corpus designed to test robustness under increasing structural complexity. We evaluate thirteen open- and closed-source systems with five complementary metrics under a unified protocol. Results show that current systems remain far from faithful transcription: performance drops sharply on complex multi-line formulas, model rankings vary across language and formula settings, and several generative models hallucinate plausible but visually unsupported corrections. OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.
CommentsCIKM 2026