发表机构
Xiaomi(小米)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Xiaomi-OCR-0,一个0.8B的统一文档解析与OCR理解模型,通过自动化数据引擎构建1.7亿样本语料库,结合渐进式训练和混合任务强化学习,在多个基准上取得领先性能。
AI 中文摘要
紧凑型OCR专用视觉语言模型在文档解析任务上表现出色,但通常依赖昂贵的监督信号,且主要聚焦于视觉-文本重建。我们推出了Xiaomi-OCR-0,一个统一的0.8B参数模型,用于文档解析和以OCR为中心的理解。我们利用自动化数据引擎构建了约1.7亿样本的OCR中心语料库,该引擎结合了专家共识、基于渲染的验证和定向合成。从Qwen3.5-0.8B出发,我们的渐进式训练方案结合了基于Q-Mask的文本锚定、持续预训练和混合任务强化学习(Mix-RL)。Xiaomi-OCR-0在Real5-OmniDocBench上达到95.24,在OmniDocBench v1.6上达到96.83,在Wild-OmniDocBench上达到87.94,同时在五个面向OCR的VQA基准上平均得分83.2。消融实验进一步表明,在充分的解析训练下,以OCR为中心的理解监督为文档解析提供了额外增益。主页:此https URL。
英文摘要
Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an average score of 83.2 across five OCR-oriented VQA benchmarks. Ablations further show that, with sufficient parsing training, OCR-centric understanding supervision provides additional gains for document parsing. Homepage: https://huggingface.co/spaces/SeerRay-Lab/Xiaomi-OCR-0.