AI 中文总结
针对文档解析语料库稀缺问题,Infinity-Parser2结合可控数据合成与多任务强化学习。贡献包括构建合成引擎及开源语料库,引入多任务奖励系统,发布Infinity-Parser2-Flash和Infinity-Parser2-Pro两个变体,后者在多项任务中达先进水平。
AI 中文摘要
我们展示了Infinity-Parser2,这是一个大型多模态模型,它将可控数据合成管道与多任务强化学习相结合,用于端到端文档解析,以解决忠实标注解析语料库长期稀缺的问题。我们有三方面贡献。首先构建了可扩展合成引擎,创建并开源了Infinity-Doc2-5M语料库。其次引入可验证的多任务奖励系统,实现跨八个协同训练目标的联合强化学习。最后发布了两个共享架构变体,Infinity-Parser2-Flash优化低延迟推理,Infinity-Parser2-Pro在关键精度设置中表现出色,在olmOCR-Bench和ParseBench上达到先进水平。
英文摘要
We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora. Our contributions are threefold. First, we build a scalable synthesis engine, pairing a controllable rendering framework with an iterative refinement loop, and use it to construct and open-source Infinity-Doc2-5M: a 5-million-sample bilingual (Chinese/English) corpus spanning diverse document types, annotated with element bounding boxes, canonical content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. Second, we introduce a verifiable, multi-task reward system that enables Joint Reinforcement Learning across eight co-trained objectives (document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding), unifying perception, structure, and reasoning in a single optimization signal. Third, we release two variants under a shared architecture: Infinity-Parser2-Flash, optimized for low-latency inference with a 3.68x throughput gain over Infinity-Parser-7B, and Infinity-Parser2-Pro, engineered for precision-critical settings. Infinity-Parser2-Pro reaches state-of-the-art 87.6% on olmOCR-Bench and 74.3% on ParseBench, surpassing DeepSeek-OCR-2, PaddleOCR-VL-1.5, and MinerU2.5, with strong generalization to charts, chemical formulas, and document VQA.