发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍拥有8亿参数的OvisOCR2文档解析模型,通过构建数据引擎,采用监督微调、强化学习、策略蒸馏和模型融合等方法训练,在多个基准测试中取得优异成绩,展现出良好的泛化性和鲁棒性。
AI 中文摘要
我们介绍了OvisOCR2,一个拥有8亿参数的文档解析模型。它被设计为端到端解析器,给定文档页面图像,能按自然阅读顺序生成Markdown表示,涵盖文本、公式、表格和视觉区域。我们构建了数据引擎,结合了经过筛选的真实文档注释与合成页面。训练方法包括监督微调、在一个拥有40亿参数分支上进行多组件奖励设计的强化学习、策略蒸馏到8亿参数模型以及模型融合。在OmniDocBench v1.6上,OvisOCR2取得了96.58的最优综合得分,在PureDocBench上也取得了75.06的最高Avg3得分。在内部基准测试中,OvisOCR2在比较方法中获得了最佳整体性能,证明了其泛化性和鲁棒性。
英文摘要
We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. We build a data engine that combines filtered real-document annotations with synthetic pages whose rendered images and Markdown targets are derived from the same HTML source. The training recipe includes supervised fine-tuning, reinforcement learning on a 4B branch with a multi-component reward design, on-policy distillation into the 0.8B model, and model fusion. On OmniDocBench v1.6, OvisOCR2 achieves a state-of-the-art overall score of 96.58, placing an end-to-end model at the top of this leaderboard previously dominated by pipeline methods and highlighting the potential of end-to-end document parsing. On PureDocBench, OvisOCR2 also achieves the highest Avg3 score of 75.06. Beyond these two public benchmarks, we evaluate OvisOCR2 on an in-house benchmark designed to cover a broader set of long-tail and challenging scenarios. OvisOCR2 obtains the best overall performance among the compared methods, providing further evidence of its generalization and robustness. OvisOCR2 is available at https://huggingface.co/ATH-MaaS/OvisOCR2.