SCAFFOLD:包含图表问答与思维链推理轨迹的大规模计算机科学研究图表结构化数据集
SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces
浏览论文内容
中文总结 AI 辅助
研究人员构建了含问答与思维链轨迹的SCAFFOLD系列计算机科学图表数据集,含三个规模子集,并用SCAFFOLD-12K在Qwen2.5-VL-3B-Instruct上完成基线实验,为视觉语言模型训练提供支撑。
中文摘要 AI 辅助
计算机科学论文严重依赖图表:架构图、系统流程图和流水线示意图,这些图表通常承载着比周围文本更多的信息。目前还没有公开的数据集将这类特定图表与其说明文字、上下文、问题、答案及逐步推理过程配对,而这正是训练视觉语言模型理解这类图表所需的内容。我们提出SCAFFOLD,这是一个包含图表问答与思维链推理轨迹的大规模计算机科学研究图表结构化数据集。该数据集由来自arXiv计算机科学论文的(图像、说明文字、上下文、问答对、思维链)元组构成,通过布局检测和PDF解析制备,并经AI辅助的问题生成步骤处理。生成的数据集包含三个规模:大规模的SCAFFOLD-157K数据集,涵盖3058篇论文、29887张图表(共157387对);中等规模的SCAFFOLD-37K数据集(36797对);小规模的SCAFFOLD-12K数据集(12000对)。我们使用SCAFFOLD-12K在Qwen2.5-VL-3B-Instruct模型上开展了基线实验。
英文摘要
Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figure with captions, context, questions, answers, and step-by-step reasoning, which is exactly what is needed to train a vision-language model to understand them. We present \textbf{SCAFFOLD}\footnote{https://github.com/theranjitraut/scaffold}, a large-scale structured dataset of computer science research figures with diagram QA and Chain-of-Thought reasoning traces. This dataset consists of (image, caption, context, question-answer, chain-of-thought) tuples from arXiv computer science papers prepared using layout detection and PDF parsing, with an AI-assisted question-generation step. The resulting large-sized SCAFFOLD-157K dataset spans 3,058 papers with 29,887 figures (157,387 pairs), a medium-sized SCAFFOLD-37K dataset (36,797 pairs), and a small-sized SCAFFOLD-12K dataset (12,000 pairs). We used SCAFFOLD-12K for baseline experiments on Qwen2.5-VL-3B-Instruct.
发表机构
- Kathmandu University(加德满都大学)
机构由 AI 辅助整理,请以论文原文为准。