发表机构
Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PolyOCR提出统一OCR基础模型系列,结合共享指令框架、大规模数据引擎和能力引导策略优化,在多个基准上达到最先进性能。
AI 中文摘要
光学字符识别(OCR)正从纯文本转录向通用视觉智能演进,要求模型在复杂视觉环境中识别、定位并推理文本信息。然而,现有OCR系统往往仅擅长部分任务,难以在多种场景下平衡识别、解析与推理。本报告中,我们提出PolyOCR——一系列不同规模的统一OCR基础模型。PolyOCR结合了共享的指令遵循框架与大规模数据引擎,将异构视觉资源转化为质量验证的OCR监督信号。我们引入能力引导的策略优化(Competence-Guided Policy Optimization),该方法将基于验证器的组相对策略优化(Group Relative Policy Optimization)与基于教师可靠性和师生能力差距的样本级路由在线策略蒸馏相结合。我们还提出OCRBench v2.1,这是对OCRBench v2的修订版,包含人工验证的标注修正和任务对齐的评分指标。在OCRBench v2.1、CC-OCR、内部KIE基准、OmniDocBench v1.6和MDPBench上的大量实验表明,PolyOCR达到了最先进或极具竞争力的性能。
英文摘要
Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision. We introduce Competence-Guided Policy Optimization, which combines verifier-based Group Relative Policy Optimization with on-policy distillation through sample-wise routing based on teacher reliability and the teacher--student competence gap. We also introduce OCRBench v2.1, our revision of OCRBench v2 with manually verified annotation corrections and task-aligned scoring metrics. Extensive experiments across OCRBench v2.1, CC-OCR, in-house KIE Benchmark, OmniDocBench v1.6 and MDPBench demonstrate that PolyOCR achieves state-of-the-art or highly competitive performance.
CommentsTechnical Report