arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.04770cs.AIcs.DB

SciIF: 科学指令遵循基准测试以实现严谨的科学智能

SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence

  • Shanghai AI Laboratory(上海人工智能实验室)
  • University of Science and Technology of China(中国科学技术大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • University of Sydney(悉尼大学)
  • Fudan University(复旦大学)
  • Tsinghua University(清华大学)
  • Beihang University(北航大学)

机构由 AI 辅助整理,请以论文原文为准。

Encheng Su, Jianyu Wu, Chen Tang, Lintao Wang, Pengze Li, Aoran Wang, Jinouwen Zhang, Yizhou Wang, Yuan Meng, Xinzhu Ma, Shixiang Tang, Houqiang Li

更新

AI总结:

SciIF是一个评估模型在科学问题解决中严格遵守约束条件能力的多学科基准测试,通过显式证据验证科学有效性,提升LLM在科学逻辑框架中的可靠性。

AI中文摘要:

随着大型语言模型(LLMs)从一般知识检索转向复杂的科学发现,其评估标准也必须纳入科学探究的严谨规范。现有基准测试存在关键盲点:通用指令遵循指标只关注表面格式,而领域特定的科学基准测试只评估最终答案的正确性,往往奖励那些通过错误原因得出正确结果的模型。为解决这一差距,我们引入了科学指令遵循:解决问题的同时严格遵守确立科学有效性的约束条件。具体而言,我们引入了SciIF,一个多学科基准测试,通过将大学级别的问题与一个固定约束目录配对,评估这一能力,涵盖三个支柱:科学条件(例如边界检查和假设)、语义稳定性(例如单位和符号惯例)以及特定过程(例如必需的数值方法)。独特的是,SciIF强调可审计性,要求模型提供显式的约束满足证据,而不是隐含的合规性。通过测量解决方案的正确性和多约束的遵守,SciIF能够对组合推理失败进行细粒度诊断,确保LLMs能够在科学的严格逻辑框架中发挥作用。

英文摘要:

As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of scientific inquiry. Existing benchmarks exhibit a critical blind spot: general instruction-following metrics focus on superficial formatting, while domain-specific scientific benchmarks assess only final-answer correctness, often rewarding models that arrive at the right result with the wrong reasons. To address this gap, we introduce scientific instruction following: the capability to solve problems while strictly adhering to the constraints that establish scientific validity. Specifically, we introduce SciIF, a multi-discipline benchmark that evaluates this capability by pairing university-level problems with a fixed catalog of constraints across three pillars: scientific conditions (e.g., boundary checks and assumptions), semantic stability (e.g., unit and symbol conventions), and specific processes(e.g., required numerical methods). Uniquely, SciIF emphasizes auditability, requiring models to provide explicit evidence of constraint satisfaction rather than implicit compliance. By measuring both solution correctness and multi-constraint adherence, SciIF enables finegrained diagnosis of compositional reasoning failures, ensuring that LLMs can function as reliable agents within the strict logical frameworks of science.

↑