OmniDocBench:具有全面标注的多类型PDF文档解析基准
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
- Shanghai AI Laboratory(上海人工智能实验室)
- Abaka AI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
OmniDocBench是一个包含九种文档来源和全面标注的基准,支持多层级评估,用于公平、多样地评测文档解析方法。
AI中文摘要:
文档内容提取是计算机视觉中的一项关键任务,支撑着大型语言模型(LLMs)和检索增强生成(RAG)系统的数据需求。尽管近期取得了进展,但由于现有基准中文档类型覆盖范围狭窄以及简化且不切实际的评估流程,当前的文档解析方法尚未得到公平和全面的评估。为弥补这些不足,我们提出了OmniDocBench,一个新颖的基准,包含来自九种文档来源的高质量标注,包括学术论文、教科书,以及更具挑战性的案例,如手写笔记和排版密集的报纸。OmniDocBench支持灵活的多层级评估——从端到端评估到基于19种布局类别和15种属性标签的任务特定及属性分析。我们对基于流程的方法和端到端视觉语言模型进行了全面评估,揭示了它们在不同文档类型上的优缺点。OmniDocBench为文档解析中公平、多样和细粒度的评估树立了新标准。数据集和代码可在https://github.com/opendatalab/OmniDocBench获取。
英文摘要:
Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have not been fairly and comprehensively evaluated due to the narrow coverage of document types and the simplified, unrealistic evaluation procedures in existing benchmarks. To address these gaps, we introduce OmniDocBench, a novel benchmark featuring high-quality annotations across nine document sources, including academic papers, textbooks, and more challenging cases such as handwritten notes and densely typeset newspapers. OmniDocBench supports flexible, multi-level evaluations--ranging from an end-to-end assessment to the task-specific and attribute--based analysis using 19 layout categories and 15 attribute labels. We conduct a thorough evaluation of both pipeline-based methods and end-to-end vision-language models, revealing their strengths and weaknesses across different document types. OmniDocBench sets a new standard for the fair, diverse, and fine-grained evaluation in document parsing. Dataset and code are available at https://github.com/opendatalab/OmniDocBench.