AI 中文总结
针对文档解析中OCR工具选择难题,尤其是标签稀缺情况,提出无标注评估框架DocOCR-Eval。它采用三阶段校正和排序策略,通过跨多个MLLM聚合改善与基于标注排名的对齐,能在现实有限标签设置中实现可靠OCR工具选择并提供实用指导。
AI 中文摘要
文档解析是诸如视觉问答和关键信息提取等文档理解任务的基础步骤,它通过提取文本、视觉和布局信息将非结构化扫描图像转换为结构化表示。尽管已经为此开发了众多光学字符识别(OCR)引擎和多模态大语言模型(MLLM),但为给定文档集选择合适的文档解析解决方案仍然具有挑战性,尤其是在标签稀缺的情况下。在这项工作中,我们在跨越不同领域和语言的多个扫描文档基准上,对各种OCR引擎和最先进的MLLM的文本识别性能进行了系统评估。由于许多OCR引擎的上下文推理能力有限且人工标注成本高,我们提出了DocOCR-Eval,这是一个用于自动OCR评估和选择的无标注评估框架。DocOCR-Eval采用三阶段校正和排序策略,在没有真实标签的情况下近似基于标注的工具排序。我们表明,跨多个MLLM聚合可逐步改善与基于标注的排名的对齐。广泛的实验进一步证明,在现实的、标签有限的设置中可以实现可靠的OCR工具选择,为跨不同真实世界文档集部署文档解析系统提供了实用指导。
英文摘要
Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information. While numerous Optical Character Recognition (OCR) engines and multimodal large language models (MLLMs) have been developed for this purpose, selecting an appropriate document parsing solution for a given document collection remains challenging, particularly in label-scarce settings. In this work, we conduct a systematic evaluation of text recognition performance across a diverse set of OCR engines and state-of-the-art MLLMs on multiple scanned document benchmarks spanning different domains and languages. Motivated by the limited contextual reasoning capabilities of many OCR engines and the high cost of manual annotations, we propose DocOCR-Eval, an annotation-free evaluation framework for automatic OCR assessment and selection. DocOCR-Eval employs a three-staged correction and ranking strategy to approximate annotation-based tool ordering without ground-truth labels. We show that aggregating across multiple MLLMs progressively improves alignment with annotation-based rankings. Extensive experiments further demonstrate that reliable OCR tool selection can be achieved in realistic, label-limited settings, providing practical guidance for deploying document parsing systems across diverse real-world document collections.
CommentsWork in progress