arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20423cs.CV

WeVisDoc:从覆盖到能力,实现稳健的端到端文档解析

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

Hao Yu, Kang Liu, Linnan Zhao, Jiabo Zhan, Chong Sun, Chen Li, Jing Lyu

首次发表
浏览论文内容

中文总结 AI 辅助

WeVisDoc提出两阶段数据中心框架,通过扩大数据覆盖并基于留出探针诊断残余错误来指导数据构建,显著提升端到端文档解析的稳健性,在多个基准上取得领先性能。

中文摘要 AI 辅助

文档解析将文档图像转换为结构化内容,并需要在不同的布局和采集条件下保持可靠的性能。然而,训练语料库偏向于常见的文档类型和干净的数字化页面,而仅扩大覆盖范围并不能指明如何解决解析器剩余的弱点。我们提出了WeVisDoc,一个两阶段的数据中心框架,用于稳健的端到端文档解析。第一阶段通过异构数据和保持结构完整的退化合成来拓宽语义、结构和外观覆盖。第二阶段使用一个留出的探针来测量第一阶段解析器在固定的视觉-结构聚类中的残余错误。这些诊断指导了针对性的数据构建和目标令牌预算的重新分配。WeVisDoc-4B在OmniDocBench v1.6上取得了95.38的总体得分,并在PureDocBench的三个赛道中平均总体得分为75.54,在所有四个设置中均排名第一,优于所比较的端到端解析器。与第一阶段相比,第二阶段在两个基准上均提高了2B和4B模型的总体得分,在退化的PureDocBench赛道上提升更大,其中4B模型在真实退化赛道上的得分提高了4.03分。

英文摘要

Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser's residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.

发表机构

  • WeChat Vision, Tencent Inc.(腾讯微信视觉)

机构由 AI 辅助整理,请以论文原文为准。

↑