发表机构
The Chinese University of Hong Kong, Shenzhen; University of Washington(香港中文大学(深圳); 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出FormStruct-Bench基准,通过分层评估表格型文档结构识别,发现现有系统在细粒度结构恢复上远弱于内容读取,为相关任务提供诊断依据。
AI 中文摘要
将表格型文档转换为机器可处理的记录,不仅需要恢复其可见内容,还需恢复组织内容的多级结构。然而,现有基准要么评估整体文档输出,要么评估常规表格网格,其聚合分数几乎无法揭示结构故障发生的位置。我们提出FormStruct-Bench,这是一个分层诊断基准,可在文档级别及逐步细化的组件级别评估表格型文档结构识别,使聚合性能可追溯至特定结构故障模式。为大规模构建可审计的真值,我们标注了70个可复用模板,并通过保留溯源的Director–Artist–Verifier流水线将其扩展为7000个验证实例;模板不相交测试集中的全部1100个实例还经过人工审核。我们的评估协议在页面、模式和组件级别使用5个主要指标和3个结构特定诊断,同时按难度、结构约束和视觉退化进行划分。在14个API托管和本地可部署系统外加2个SFT变体中,最佳文档级别分数达到83.85%,而最佳报告的细粒度结构分数仍低于18%。这些结果揭示了读取文档内容与恢复可靠表格型理解所需的层次结构和区域组织之间存在显著差距。
英文摘要
Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure that organizes it. However, existing benchmarks evaluate either holistic document outputs or conventional table grids, and their aggregate scores provide little insight into where structural failures occur. We introduce FormStruct-Bench, a hierarchical and diagnostic benchmark that evaluates table-form document structure recognition at both the document level and progressively finer component levels, allowing aggregate performance to be traced back to specific structural failure modes. To construct auditable ground truth at scale, we annotate 70 reusable templates and expand them into 7,000 verified instances through a provenance-preserving Director--Artist--Verifier pipeline; all 1,100 instances in the template-disjoint test set additionally receive human review. Our evaluation protocol uses five primary metrics and three structure-specific diagnostics across page, schema, and component levels, together with slices over difficulty, structural constraints, and visual degradation. Across 14 API-hosted and locally deployable systems plus two SFT variants, the best document-level score reaches 83.85%, whereas the best reported fine-grained structural score remains below 18%. These results reveal a pronounced gap between reading document content and recovering the hierarchy and regional organization required for reliable table-form understanding.