arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11192cs.CV

GDP.pdf:针对专业PDF文档的基础多模态推理基准测试

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

Suhaas Garre, Emily Ritchie, Sushant Mehta, Edwin Chen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对专业PDF文档构建多模态推理基准测试GDP.pdf,由专业人员编写问题-文档对,通过严格筛选保留问题,有详细评分标准和能力分类。评估七个前沿模型,发现多数错误源于特定模式,公开了完整基准测试。

中文摘要 AI 辅助

在专业领域中,大量日常工作发生在PDF文件内,如福利包、租约、数据表、临床指南、施工计划等。文档AI的基准测试通常孤立地衡量所需能力,如OCR、布局分析等。而这个基准测试旨在直接衡量模型能否回答专业领域人员关于特定PDF的实际问题。它由十个领域的专业人员编写的问题-文档对组成,只有当至少两个前沿多模态模型以重要方式答错时,候选问题才会保留。每个项目都有原子标准的评分标准,可报告分级评分和严格的任务级通过率,且每个项目都根据三个层次的十一种能力分类进行标记。对七个前沿模型在100个项目的基准测试上进行评估,最佳模型仅通过15%的项目,最差的通过1%。大多数错误可追溯到一小部分反复出现的损失模式,如表格对齐错误、图表误读等。完整的100个项目基准测试可在指定网址公开获取。

英文摘要

A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, construction plans. Benchmarks for document AI have generally measured the required capabilities in isolation: OCR, layout analysis, chart reasoning, table QA, document VQA. A high score on any one of them does not necessarily reveal whether a model can answer a realistic question that someone in the field would actually ask about a specific PDF. GDP_pdf is a benchmark built to measure this directly. It consists of question-document pairs authored by working professionals in ten fields, and a candidate question was kept only when at least two frontier multimodal models failed it in a way that mattered: a wrong answer, missed decisive evidence, or a fabricated claim, rather than a superficial difference such as style. Each item comes with a rubric of atomic criteria, so we can report a graded rubric score as well as a strict task-level pass rate, and each item is tagged against a taxonomy of eleven capabilities in three tiers, spanning text extraction and grounding, table and chart comprehension, cross-referencing, spatial reasoning, and abstention on unsupported queries. We report results for seventeen frontier models on the 100-item benchmark: the best model passes only 30.7% of the items and the worst passes 2%. Most errors trace back to a small set of recurring loss patterns: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text.

发表机构

  • Surge AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑