从诊断到修正:真实世界表格解析的基准测试与改进
From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing
浏览论文内容
中文总结 AI 辅助
本文针对真实世界表格解析的缺陷,构建诊断基准TableParseMap,提出DEC框架,无需重训即可提升冻结解析器性能,在TableParseMap上TEDS获显著提升。
中文摘要 AI 辅助
近期,文档解析器在OmniDocBench v1.6上的表格TEDS分数达到93以上,但社区反馈及本文审计显示,其在复杂真实世界表格上仍存在持续的失败。为量化这一差距,本文引入TableParseMap,这是一个包含916张真实表格的诊断基准,分为5种具有挑战性的场景和9种失败类型。经评估,最强的解析器仅达到85.03的TEDS分数,表明聚合基准分数掩盖了大量缺陷。本文分析将这些失败归因于三个互补的局限性:大型表格超出单次处理的可靠规模,薄弱或模糊的视觉线索阻碍结构感知,重建的表格可能与图像视觉不一致。因此,本文提出DEC(Decompose--Enhance--Correct,分解-增强-修正),这是一种视觉一致性引导的智能体框架,可改进冻结的表格解析器且无需重新训练。DEC使用通用VLM作为控制器:分解(Decompose)沿结构感知边界划分大型表格,增强(Enhance)暴露薄弱的视觉证据并重新解析变换后的视图,修正(Correct)诊断并修复残留错误。视觉一致性门(VC-Gate)选择性触发干预,视觉一致性排序器(VC-Ranker)验证候选更新并支持回滚,推理时无需真值HTML。本文还通过离线指标和跨模型共识,从4556个候选中得到包含1977张表格的Consensus-Hard Set。在三个冻结解析器上,DEC平均提升TEDS 1.57个点;在TableParseMap上,整体提升达1.89个点,结构错误提升2.62个点,大型表格提升5.66个点。
英文摘要
Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To quantify this gap, we introduce TableParseMap, a diagnostic benchmark of 916 real-world tables organized into five challenging scenarios and nine failure types. The strongest evaluated parser achieves only 85.03 TEDS, showing that aggregate benchmark scores conceal substantial weaknesses. Our analysis attributes these failures to three complementary limitations: large tables exceed the reliable processing scale of a single pass, weak or ambiguous visual cues hinder structure perception, and the reconstructed table may remain visually inconsistent with the image. We therefore propose DEC (Decompose--Enhance--Correct), a visual-consistency-guided agentic framework that improves frozen table parsers without retraining. DEC uses a general VLM as the controller: Decompose partitions large tables along structure-aware boundaries, Enhance exposes weak visual evidence and reparses transformed views, and Correct diagnoses and repairs residual errors. A Visual Consistency Gate (VC-Gate) selectively triggers intervention, while a Visual Consistency Ranker (VC-Ranker) verifies candidate updates and supports rollback without ground-truth HTML at inference time. We further derive a 1,977-table Consensus-Hard Set from 4,556 candidates through offline metrics and cross-model consensus. Across three frozen parsers, DEC improves TEDS by 1.57 points on average; on TableParseMap, gains reach 1.89 points overall, 2.62 on structural errors, and 5.66 on large tables.
发表机构
- Zhejiang University(浙江大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。