arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30214cs.AI

SPARK:基于大规模科学文献的骨架引导推理合成

SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature

  • Shanghai AI Laboratory(上海人工智能实验室)
  • University of Science and Technology of China(中国科学技术大学)
  • East China Normal University(华东师范大学)
  • Peking University(北京大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

Yu Li, Wei Li, Xin Gao, Mengyuan Sun, Xiaoyang Wang, Qizhi Pei, Lijun Wu

AI总结:

针对开源模型科学推理数据不足的问题,提出基于Sci-Base的SPARK框架构建Spark-234K科学推理数据集,该数据集难度与多样性更优,仅需更少训练样本即可实现更强性能。

AI中文摘要:

开源模型在科学推理方面仍面临挑战,主要原因是缺乏高质量的科学推理数据。现有数据集通常以事实回忆或公式化问题解决为主,对机制理解、证据支撑推理和假设评估的重视有限。为解决这一问题,我们引入SPARK(Scientific Paper Abstracted Reasoning sKeleton,即科学论文抽象推理骨架),这是一个基于Sci-Base的面向论文的合成框架,Sci-Base是涵盖10个科学学科的大规模研究论文语料库。SPARK不直接将论文转换为问答对,而是将论文的主张-证据-推导结构作为推理合成的基本单元。具体而言,SPARK(1)将每篇论文提炼为紧凑的推理骨架,捕捉其核心主张和支撑证据,从而实现自包含的问题生成;(2)从四个科学视角合成推理任务:机制推理、假设证伪、定量推导和边界校准。最后,一致性验证阶段会进一步移除无支撑或矛盾的输出。利用该框架,我们构建了Spark-234K,这是一个科学推理数据集,其难度和多样性显著高于现有资源。实验表明,Spark-234K在性能上始终优于现有科学推理数据集,且使用少得多的训练样本即可实现更强的表现。

英文摘要:

Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited emphasis on mechanism understanding, evidence-grounded reasoning, and hypothesis evaluation. To address this, we introduce SPARK (Scientific Paper Abstracted Reasoning sKeleton), a paper-oriented synthesis framework built on Sci-Base, a large-scale corpus of research papers spanning 10 scientific disciplines. Instead of directly converting papers into question-answer pairs, SPARK treats the claim-evidence-derivation structure of a paper as the fundamental unit of reasoning synthesis. Specifically, SPARK (1) distills each paper into a compact reasoning skeleton capturing its central claims and supporting evidence, enabling self-contained question generation, and (2) synthesizes reasoning tasks from four scientific perspectives: mechanistic reasoning, hypothesis falsification, quantitative derivation, and boundary calibration. A final consistency verification stage further removes unsupported or contradictory outputs. Using this framework, we construct Spark-234K, a scientific reasoning dataset with substantially higher difficulty and diversity than existing resources. Experiments show that Spark-234K consistently outperforms existing scientific reasoning datasets while achieving stronger performance with significantly fewer training samples.

补充信息

↑