arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Jina-OCR-v1:结合推测解码与密集可验证奖励的高效文档解析模型

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

Alejandro Barón García, Feng Wang, Emilia Garcia Casademont, Han Xiao

arXiv 2609.03181首次发表:更新:

发表机构

Jina AI(金纳人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Jina-OCR-v1是专为低预算GPU设计的端到端文档解析模型,结合DeepSeek-OCR的编码器与解码器,采用FastMTP推测解码和密集可验证奖励训练,在多个基准上取得优异性能且解码速度提升显著。

AI 中文摘要

我们提出Jina-OCR-v1,一款专为低预算GPU设计的端到端文档解析模型。它结合了DeepSeek-OCR的压缩视觉编码器与3B混合专家解码器(每token激活约5.7亿参数),以及FastMTP推测解码头,该头在K=3个预测步骤中递归共享单个草稿块。贪心验证使解码无损失。后训练结合了指令对齐、针对困难文档的鲁棒性微调,以及在密集可验证奖励下的GRPO:确定性公式、表格和结构检查,可授予部分分数。训练数据混合了清洗后的公开语料库与针对性合成页面。在默认动态分辨率设置下,Jina-OCR-v1在OmniDocBench v1.6上得分91.14,在olmOCR-Bench上得分83.4,且在我们的对比中达到最高页面吞吐量,为2.57页/秒。在NVIDIA L4等低预算GPU上,FastMTP使解码速度较贪心自回归解码翻倍。该模型可在指定URL公开获取。

英文摘要

We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and reaches the highest page throughput in our comparison at 2.57 pages per second. On a low-budget GPU such as the NVIDIA L4, FastMTP doubles decoding speed over greedy autoregressive decoding. The model is publicly available at https://huggingface.co/jinaai/jina-ocr-v1.

Comments15 pages, 5 figures, 8 tables. Model at https://huggingface.co/jinaai/jina-ocr-v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑