SciDocBench:面向科学文档理解的以工作流为中心的基准与数据管道
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
- The Chinese University of Hong Kong(香港中文大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- New York University(纽约大学)
- Fudan University(复旦大学)
- Shanghai Jiao Tong University(上海交通大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- JD Explore Academy(京东探索研究院)
- Centre for Perceptual and Interactive Intelligence (CPII) Limited(感知与交互智能中心(CPII)有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出以工作流为中心的科学文档理解基准SciDocBench,结合多模态评估与证据图表示构建数据集,发现现有模型性能有限,形成评估到训练框架以改进科学文档助手。
AI中文摘要:
科学论文要求模型在推理时需同时处理文本、公式、图表、表格、代码和数据集,同时保留支撑证据的来源。现有基准通常孤立地评估这些能力,无法明确多模态模型是否能支持真实的科学阅读工作流。本文提出SciDocBench,这是一个以工作流为中心的科学文档理解基准,包含124道经专家编写和难度筛选的问题,分为7个研究助手能力组和5个科学领域的19个子任务。每个问题在4种匹配条件下实例化,结合英文或中文问题、全图像优先或交错文档表示,共产生496个评估实例用于受控分析。评估中最强的系统仅达到62.6/100,在文档感知、证据锚定、验证和跨文档推理方面存在明显弱点。为将这些诊断转化为可扩展的训练信号,本文提出SciDocIR,一种保留科学文档对象、布局与交叉引用关系及来源的类型化证据图表示。基于SciDocIR构建SciDocDataset,包含约15000个监督微调样本和8000个强化学习样本,覆盖14个可验证子任务。SciDocBench、SciDocIR和SciDocDataset共同构成用于诊断和改进科学文档助手的评估到训练框架,项目页面可通过该https URL访问。
英文摘要:
Scientific papers require models to integrate evidence across text, equations, figures, tables, code, and datasets while preserving its provenance. Beyond answer correctness, scientific reading requires verifiable outputs from operations such as evidence localization, definition extraction, and consistency checking. We introduce SciDocBench, a workflow-centered benchmark targeting these operations through 124 expert-authored and difficulty-screened questions across seven research-assistant capability groups, 19 subtasks, and five scientific domains. Each question is instantiated in four matched settings formed by pairing its bilingual variants with the All Images First and Markdown Interleaved document representations, yielding 496 evaluation instances. The strongest evaluated model, Claude-Opus-5, scores 62.6 out of 100, with remaining gaps in evidence localization, structured information extraction, cross-document synthesis, and robustness to document representation. To convert these diagnostics into scalable training signals, we introduce SciDocIR, a structured representation of scientific document objects, layout and cross-reference relations, and provenance. Using SciDocIR, we construct SciDocDataset, which contains 4K supervised fine-tuning instances and 10K reinforcement-learning instances, built on 14 verifiable training subtasks. Post-training Qwen3.6-27B on task-aligned data improves its SciDocBench score from 40.03 to 45.33 with supervised fine-tuning and to 45.74 with subsequent reinforcement learning. Both adapted models preserve DocVQA and InfoVQA performance and improve ChartQA accuracy over the original model by 0.80 and 3.40 points, respectively. Together, SciDocBench, SciDocIR, and SciDocDataset connect capability diagnosis with verifiable training-data construction for scientific-document assistants.