AI 中文总结
LongDocBench是含85份长文档的基准,可评估目录层级与上下文关系恢复任务,实验显示该基准结构能提升长文档问答推理,相关资源已公开。
AI 中文摘要
将视觉文档解析为机器可读的表示形式是文档智能的基础。现有基准侧重于页面级元素识别、阅读顺序、公式识别和表格结构,然而长文档还需要文档级结构恢复,包括重建跨页目录(TOC)层级,以及识别表格、图表与其标题、注释、来源之间的类型化链接,这类链接常为一对多形式。由于这些结构仅被部分覆盖或包含在更广泛的解析协议中,现有基准无法直接评估两个关键文档级任务:目录层级恢复和上下文关系恢复。为对这两个任务进行基准测试,我们推出LongDocBench,包含85份真实世界的财务报告、教科书和学术论文,共2582页,单份文档最多105页。它提供经人工验证的注释,涵盖3937个标题节点(平均节点深度3.55,最大深度9),以及2680个表格和图表对象上注释的3258个上下文关系。我们进一步评估这些结构的下游效用和可恢复性。长文档问答实验表明,经人工验证的目录层级和上下文关系可提升推理能力,二者结合能提供互补益处。同时,代表性文档解析器虽在页面级任务上表现强劲,但在这两个恢复任务上仍存在局限。为支持进一步研究,我们公开发布LongDocBench及其评估协议和可复现测试平台,以推进长文档的文档级结构恢复研究。
英文摘要
Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and figures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing benchmarks cannot directly evaluate two key document-level tasks: \emph{Table-of-Contents Hierarchy Recovery} and \emph{Contextual Relationship Recovery}. To benchmark these two tasks, we introduce \textsc{LongDocBench}, comprising 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships annotated across 2,680 table and figure objects. We further evaluate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits. Meanwhile, representative document parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release \textsc{LongDocBench} and its evaluation protocol and reproducible testbed for advancing document-level structure recovery in long documents.
Commentspreprint, under review