发表机构
Glean Technologies, Inc(Glean技术公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多页文档结构分析难题,提出树引导自回归框架DOSA,逐块处理文档,融合多特征预测关系来构建语义树,实验表明该方法在DocHieNet等基准上效果显著,提升了F1分数和TEDS分数。
AI 中文摘要
在视觉丰富的文档中,信息不仅编码在表格、标题和文本块等单个页面对象中,还存在于它们之间的结构关系中,因此文档结构分析对信息检索和文档理解至关重要。然而,在具有长距离依赖和异构布局的多页文档中准确推断此类关系仍具有挑战性。为解决此问题,我们提出了一个树引导自回归框架,称为文档结构分析器(DOSA),用于推断页面对象之间的关系并重建文档级语义树。DOSA逐块处理文档,融合每个页面对象的视觉、文本和布局特征,并预测层次和顺序关系。预测的关系用于增量构建语义树,然后将其用作结构上下文来指导后续块的推断。在五个基准上的实验结果证明了DOSA的有效性,在最具挑战性的多页层次基准DocHieNet上,F1分数提高了4分,TEDS分数提高了19分。
英文摘要
In visually-rich documents, information is encoded not only in individual page objects such as tables, headers, and text blocks, but also in the structural relations among them, making document structure analysis fundamental to information retrieval and document understanding. However, accurately inferring such relations remains challenging in multi-page documents with long-range dependencies and heterogeneous layouts. To address this, we propose a tree-guided and self-regressive framework, termed DOcument Structure Analyzer (DOSA), for inferring relations among page objects and reconstructing document-level semantic trees. DOSA processes documents chunk-by-chunk, fusing visual, textual, and layout features for each page object and predicting hierarchical and ordering relations. The predicted relations are used to incrementally construct a semantic tree, which is then leveraged as structural context to guide inference on subsequent chunks. Experimental results on five benchmarks demonstrate the effectiveness of DOSA, with improvements of up to 4 F1 points and 19 TEDS points on DocHieNet, the most challenging multi-page hierarchy benchmark.