arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

崩溃,而非复杂性:面向端到端文档解析的失败条件分解修复

Collapse, Not Complexity: Failure-Conditioned Decomposition Repair for End-to-End Document Parsing

Xingyu Lin, Dehui Du

arXiv 2609.23592首次发表:更新:

发表机构

Software Engineering Institute, East China Normal University(华东师范大学软件工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现文档解析失败源于页面崩溃而非复杂性,提出基于失败条件的分解修复方法,在低令牌消耗下显著提升解析质量。

AI 中文摘要

端到端文档解析器日益为复杂页面提供可选的推理模式。在一个包含180页、按熵分层发现的样本上,使用一个冻结的4B检查点,复杂性是错误的决策变量。推理将平均质量降低2.21 Overall,同时消耗1.54倍的令牌;一个预注册的仅输入模型无法预测其符号化收益(留出法AUROC为0.47,与随机猜测无异)。收益集中在普通通过(ordinary pass)已经崩溃的页面上,而这些页面看起来并不复杂:共享崩溃的布局熵低于健康页面,但消耗的令牌却是后者的19倍,表现为退化性重复,即使将预算加倍也无法解决。切换模式很少能修复这些页面:83%的页面在推理下仍会复发。我们转而从普通通过的轨迹中检测崩溃,通过投影对页面进行分解,并重新解析每个区域。修复在1.13倍令牌消耗下获得1.40 Overall的提升(95%置信区间[0.68, 2.16]),在三个检查点上均可复现,并且在所有参数冻结的情况下,在剩余的1,175个基准页面上获得2.41的提升(置信区间[1.64, 3.46])。

英文摘要

End-to-end document parsers increasingly offer an optional reasoning mode for complex pages. On a 180-page entropy-stratified discovery sample with one frozen 4B checkpoint, complexity is the wrong decision variable. Reasoning lowers mean quality by 2.21 Overall at 1.54x tokens; a preregistered input-only model cannot predict its signed benefit (held-out AUROC 0.47, indistinguishable from chance). The benefit concentrates on pages whose ordinary pass has already collapsed, and they do not look complex: shared collapses have lower layout entropy than healthy ones yet consume 19x the tokens as degenerate repetition that doubling the budget does not cure. Switching modes rarely repairs them: 83% recur under reasoning. We instead detect collapse from the ordinary-pass trace, decompose the page by projection, and re-parse each region. Repair gains 1.40 Overall (95% CI [0.68, 2.16]) at 1.13x tokens, replicates across three checkpoints, and, with all parameters frozen, gains 2.41 (CI [1.64, 3.46]) on the remaining 1,175 benchmark pages.

Comments5 pages, 3 figures, 2 tables. Submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑