发表机构
ScienciaLAB; Helmholtz-Zentrum Hereon(ScienciaLAB; 亥姆霍兹-盖斯特哈赫特研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对GROBID提出布局引导掩码方法,用轻量级CPU检测器定位图形表格等区域并路由或丢弃令牌,在多个语料库上提升结构解析性能,且成本远低于GPU系统。
AI 中文摘要
将学术PDF转换为机器可读的全文仍然是大型信息系统的一个瓶颈。最近的基于视觉的解析器提高了准确性,但需要GPU,并且可能向提取的文本中引入噪声。GROBID是一种在CPU上运行的模块化字体流解析器,是结构化科学文章的事实标准,并支撑着几个最大的开放学术语料库。我们将其与一个轻量级CPU检测器配对,该检测器定位图形、表格和旁文本(页眉、页脚、页码)区域,编码为类型化区域掩码,其令牌被路由到GROBID的专门模型或被丢弃。在两个PMC语料库上,即Bioinformatics(1,926篇文章)和Materials Science(2,595篇),使用带有章节感知结构协议对照JATS进行评分,我们的扩展在大多数指标上优于普通GROBID(NS $+0.025$/$+0.013$;Materials Science上段落召回率$+0.086$,$d_z{=}1.08$),并且标题链接的图形恢复在两个语料库上都有所改善。在外部Table-BRGM基准上,表格检测恢复F1从$0.16$到$0.94$,表格结构随之改善(GriTS-Top从$0.27$到$0.78$,低于最强的GPU系统)。在正文文本上,与四个基于视觉的系统(Docling、MinerU、olmOCR、this http URL)相比,它在两个语料库上具有最佳的段落精确率,在Materials Science上具有最佳的章节检测,并且字符错误率与最佳GPU解析器相差在0.004以内。在CPU上端到端运行,其成本比最便宜的GPU系统(Docling)低$2.7$--$3.2\times$,比生成式解析器低$10$--$14\times$。
英文摘要
Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems. Recent vision-based parsers improve accuracy, but need GPUs and may introduce noise into the extracted text. GROBID, a modular font-stream parser running on CPU, is the de-facto standard for structuring scientific articles and underpins several of the largest open scholarly corpora. We pair it with a lightweight CPU detector localising figure, table, and paratext (header, footer, page number) regions, encoded as typed-area masks whose tokens are routed to GROBID's specialised models or discarded. On two PMC corpora, Bioinformatics (1,926 articles) and Materials Science (2,595), scored against JATS with a section-aware structural protocol, our extension improves over plain GROBID on most metrics (NS $+0.025$/$+0.013$; $+0.086$ paragraph recall on Materials Science, $d_z{=}1.08$), and caption-linked figure recovery improves on both corpora. On the external Table-BRGM benchmark, table detection recovers F1 $0.16 \to 0.94$ and table structure follows (GriTS-Top $0.27 \to 0.78$, below the strongest GPU system). On body text, against four vision-based systems (Docling, MinerU, olmOCR, dots.ocr), it has the best paragraph precision on both corpora, the best section detection on Materials Science, and a character error rate within 0.004 of the best GPU parser. End-to-end on CPU, it costs $2.7$--$3.2\times$ less than the cheapest GPU system (Docling) and $10$--$14\times$ less than generative parsers.