OCR-EDR:面向闭环OCR改进的渲染感知诊断与修复
OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement
浏览论文内容
中文总结 AI 辅助
提出OCR-EDR框架,构建OCRErrBench数据集与DocEDR模型,实现OCR错误的细粒度诊断与迭代修复,提升OCR在公式等场景的性能。
中文摘要 AI 辅助
尽管文档OCR系统在常规文档上的表现日益出色,但复杂公式、结构化文本和长尾格式仍易出错。OCR预测可能会遗漏细粒度内容或生成无依据的输出,同时必须处理同一可见内容的等效编码。现有OCR评估方法大多报告聚合指标,对分析案例级错误和提升OCR性能的支持有限。我们提出OCR-EDR(OCR错误诊断与修复),这是一种从细粒度诊断推进到迭代修复的渲染感知框架。给定源图像、可编辑OCR预测及其渲染图像,OCR-EDR首先联合评估预测及其渲染是否与源图像一致,保留有效预测(包括渲染等效预测),同时诊断并定位真实错误;随后应用可执行编辑,并可请求更新后的渲染图像以进行迭代重新评估。我们从多样的真实OCR预测中构建OCRErrBench,涵盖文本与公式、精确及渲染等效正例、真实错误,并开发DocEDR模型以执行诊断-修复循环。在OCRErrBench上,DocEDR达到94.78%的诊断准确率;它将86.23%的错误输入修复至视觉一致性,在DOCRcaseBench上较DOCR-Inspector-7B将公式Case-F1提升30.99个百分点,在UniMER-Test的四个OCR系统识别的Bad子集上将公式CDM提升多达4.62个百分点。这些结果表明,OCR-EDR将细粒度OCR分析转化为可验证的修正和性能提升。
英文摘要
Although document OCR systems perform increasingly well on routine documents, complex formulas, structured text, and long-tail formats remain error-prone. OCR predictions may omit fine-grained content or hallucinate unsupported outputs, while equivalent encodings of the same visible content must be accommodated. Existing OCR evaluation methods mostly report aggregate metrics, offering limited support for analyzing case-level errors and improving OCR performance. We propose OCR-EDR (OCR Error Diagnosis and Repair), a rendering-aware framework that advances from fine-grained diagnosis to iterative repair. Given a source image, an editable OCR prediction, and its rendered image, OCR-EDR first jointly assesses whether the prediction and its rendering are consistent with the source, preserving valid predictions, including rendering-equivalent ones, while diagnosing and localizing genuine errors. It then applies executable edits and may request an updated rendering for iterative reassessment. We construct OCRErrBench from diverse real OCR predictions, covering text and formulas, exact and rendering-equivalent positives, and genuine errors, and develop the DocEDR model to execute the diagnosis--repair loop. On OCRErrBench, DocEDR achieves 94.78% diagnostic accuracy. It repairs 86.23% of erroneous inputs to visual consistency, raises formula Case-F1 by 30.99 percentage points over DOCR-Inspector-7B on DOCRcaseBench, and improves formula CDM by up to 4.62 percentage points on the identified Bad subsets of four OCR systems on UniMER-Test. These results show that OCR-EDR turns fine-grained OCR analysis into verified corrections and performance gains.
发表机构
- WeChat Vision, Tencent(腾讯微信视觉)
机构由 AI 辅助整理,请以论文原文为准。