扩散草稿,自回归验证:利用自推测解码加速文档OCR
Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding
浏览论文内容
中文总结 AI 辅助
针对自回归OCR推理慢的问题,提出参数共享的GravityOCR模型,结合扩散并行草稿与自回归验证,在OmniDocBench上提升精度并实现显著加速。
中文摘要 AI 辅助
自回归OCR视觉语言模型能准确地将文档图像转换为文本和结构化标记,但每个输出token需要一个顺序解码步骤,限制了推理速度。与开放式文本生成不同,OCR输出与输入图像强相关,这使得基于扩散的并行生成具有前景。然而,当在一步扩散中预测多个token时,每个token都在其他token已知之前被预测,直接提交它们可能会引入错误。因此,我们提出了GravityOCR,一个参数共享的AR块扩散模型,联合训练用于并行草稿和因果AR验证。在提交前验证草稿使模型每轮能提交多个输出token,而无需单独的草稿网络。因果AR路径还使得GRPO能够使用序列级和结构级OCR奖励,避免了扩散轨迹似然估计,同时更新共享的草稿器参数。在OmniDocBench v1.6上,AR路径GRPO将总体得分从94.92提升至95.16,且不降低扩散草稿效率,最终模型仍接近原始GLM-OCR的95.48分。在SGLang服务部署中,GravityOCR每次前向传递平均提交9.7个输出token,在区域裁剪上实现了3.94倍的仅解码加速,在端到端页面处理上比AR解码实现了1.32倍的加速。
英文摘要
Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed. Unlike open-ended text generation, OCR outputs are strongly grounded in the input image, making diffusion-based parallel generation promising. However, when several tokens are predicted in one diffusion step, each is predicted before the others are known. Committing them directly can therefore introduce errors. We therefore introduce GravityOCR, a parameter-shared AR-block-diffusion model jointly trained for parallel drafting and causal AR verification. Verifying drafts before commitment lets the model commit multiple output tokens per round without a separate drafting network. The causal AR path also enables GRPO with sequence- and structure-level OCR rewards, avoiding diffusion-trajectory likelihood estimation while updating the shared drafter parameters. On OmniDocBench v1.6, AR-path GRPO improves the Overall score from 94.92 to 95.16 without reducing diffusion drafting efficiency, while the final model remains close to the original GLM-OCR score of 95.48. In an SGLang serving deployment, GravityOCR commits an average of 9.7 output tokens per forward pass and achieves a $3.94\times$ decode-only speedup on region crops and a $1.32\times$ end-to-end page-processing speedup over AR decoding.
发表机构
- Trillion Labs
- Korea University(高丽大学)
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。