发表机构
Guangdong OPPO Mobile Telecommunications Corp.,Ltd; Adelaide University(广东欧珀移动通信有限公司; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SP-DocReader,一种针对监督微调后OCR残留错误的自博弈框架,通过阅读差异掩码与聚焦保真度损失训练OCR模块,在Qwen3-VL-4B等基准上显著降低字符错误率并提升DocVQA的ANLS指标。
AI 中文摘要
在有限的输入和训练预算下,视觉语言模型仍难以实现准确的页面转录。本文提出SP-DocReader,一种针对监督微调(SFT)后残留错误的光学字符识别(OCR)自博弈框架。其中,阅读差异掩码(Reading Discrepancy Masking)通过最长公共子序列对齐参考文本与模型生成的token,再利用完整条件前缀对未匹配位置打分;聚焦保真度损失(Focused Fidelity Loss)则在未匹配的真实位置添加直接负对数似然监督。仅训练OCR模块,主干网络保持冻结。我们推导了组合梯度以区分相对分数优化与直接监督。与SFT-2相比,SP-DR-3在两个主干网络上均降低了Vary-600K的字符错误率;在Qwen3-VL-4B上,其字符错误率降低约54%,DocVQA的平均归一化莱文斯坦相似度(ANLS)提升3.7个百分点。这些结果表明,将自博弈训练聚焦于监督微调后残留的差异具有重要价值。
英文摘要
Accurate page transcription remains difficult for vision language models under limited input and training budgets. We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning. Reading Discrepancy Masking aligns reference and generated model tokens through a longest common subsequence, then scores unmatched positions with their full conditioning prefixes. Focused Fidelity Loss adds direct negative log-likelihood supervision at unmatched ground-truth positions. Only the OCR module is trained, while the backbone remains frozen. We derive the combined gradient to distinguish relative score optimization from direct supervision. Compared with SFT-2, SP-DR-3 reduces Vary-600K character error rate on both backbones. On Qwen3-VL-4B, it reduces character error rate by approximately 54 percent and improves DocVQA Average Normalized Levenshtein Similarity (ANLS) by 3.7 points. These results show the value of focusing self-play training on the discrepancies that remain after supervised fine-tuning.