arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Infinity-Parser2技术报告

Infinity-Parser2 Technical Report

Zuming Huang, Jun Huang, Kexuan Ren, Baode Wang, Weizhen Li, Jianming Feng, Yu Wang, Yichen Yao, Shijun Lin, Yige Tang, Cheng Peng, Weidi Xu, Wei Chu, Yinghui Xu, Yuan Qi

arXiv 2607.07836首次发表:更新:

AI 中文总结

针对文档解析语料库稀缺问题,Infinity-Parser2结合可控数据合成与多任务强化学习。贡献包括构建合成引擎及开源语料库,引入多任务奖励系统,发布Infinity-Parser2-Flash和Infinity-Parser2-Pro两个变体,后者在多项任务中达先进水平。

AI 中文摘要

我们展示了Infinity-Parser2,这是一个大型多模态模型,它将可控数据合成管道与多任务强化学习相结合,用于端到端文档解析,以解决忠实标注解析语料库长期稀缺的问题。我们有三方面贡献。首先构建了可扩展合成引擎,创建并开源了Infinity-Doc2-5M语料库。其次引入可验证的多任务奖励系统,实现跨八个协同训练目标的联合强化学习。最后发布了两个共享架构变体,Infinity-Parser2-Flash优化低延迟推理,Infinity-Parser2-Pro在关键精度设置中表现出色,在olmOCR-Bench和ParseBench上达到先进水平。

英文摘要

We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora. Our contributions are threefold. First, we build a scalable synthesis engine, pairing a controllable rendering framework with an iterative refinement loop, and use it to construct and open-source Infinity-Doc2-5M: a 5-million-sample bilingual (Chinese/English) corpus spanning diverse document types, annotated with element bounding boxes, canonical content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. Second, we introduce a verifiable, multi-task reward system that enables Joint Reinforcement Learning across eight co-trained objectives (document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding), unifying perception, structure, and reasoning in a single optimization signal. Third, we release two variants under a shared architecture: Infinity-Parser2-Flash, optimized for low-latency inference with a 3.68x throughput gain over Infinity-Parser-7B, and Infinity-Parser2-Pro, engineered for precision-critical settings. Infinity-Parser2-Pro reaches state-of-the-art 87.6% on olmOCR-Bench and 74.3% on ParseBench, surpassing DeepSeek-OCR-2, PaddleOCR-VL-1.5, and MinerU2.5, with strong generalization to charts, chemical formulas, and document VQA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑