发表机构
Huazhong University of Science and Technology; Kingsoft Office(华中科技大学; 金山办公软件)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对主流视觉编码器难以应用于文档图像的问题,提出MonkeyOCRv2模型。通过构建大型文档图像预训练语料库及采用联合预训练策略,在多个文档分析任务中提升性能,并验证其作为视觉编码器在文档解析和理解任务中的有效性。
AI 中文摘要
主流视觉编码器在自然图像上预训练,因文档图像中密集文本和精细字符笔画需字符级视觉感知,无法直接应用于文档图像。本文提出用于文档人工智能的视觉-文本预训练模型MonkeyOCRv2。首先构建了包含11300万张17种语言图像的MonkeyDoc v2文档图像预训练语料库。其次提出联合学习图像到文本生成和像素级文档重建的预训练策略。在五个文档分析任务上进行大量实验,结果表明MonkeyOCRv2在所有任务中均提升了性能,还验证了其作为多模态大语言模型视觉编码器在文档解析和理解任务中的有效性。
英文摘要
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11$\times$ smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.