arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

专家验证的STEM问答

Expert-validated STEM QA

Kihwan Han, Saurabh Patil, Chinmayee Shukla, Abhinav Sharma, Marko Pavlovic, Anshuman Lall, Mahesh Joshi

arXiv 2608.28591首次发表:更新:

发表机构

Turing(图灵公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出由241名领域专家创建的专家验证STEM问答数据集,解决现有STEM数据集的不足,该数据集可作为基准评估前沿AI模型,还能提升开源模型性能,已开源部分数据。

AI 中文摘要

人工智能的最新进展正帮助科学家在数学、医学、材料科学等领域取得突破,面向AI模型的新型评估数据集推动了这类AI领域的进步。在STEM领域,前沿模型已消耗了大部分可用的在线数据,因此需要人类创建的数据集来整理该领域顶尖专家的知识。目前已有多个STEM数据集供该领域研究人员使用,但这些数据集存在一些不足,有待改进,具体包括:(1)模型在这些数据集上的性能已趋于饱和,无法开展有意义的评估;(2)分类分布不均衡;(3)多项选择题的形式与科学家在现实世界中使用AI的方式不匹配;(4)部分答案和推理依据不准确,这一问题部分源于竞赛式的数据收集方式和限时审核流程。本研究推出了“专家验证的STEM问答”(Expert-validated STEM QA),这是一个由241名领域专家创建的、涵盖物理、化学、生物和数学领域的高质量专家验证STEM数据集,样本量N=398。我们完成了四项工作:(1)精心设计了分类体系,确保分布均衡;(2)采用质量驱动的激励机制审核出题者;(3)开展多轮审核,由领域专家基于共识验证修订内容;(4)将数据集创建为可验证的问答格式。研究显示,前沿AI模型在该数据集上作为基准的性能较低,不足25%;在独立的私有版本数据集(样本量N=2000)上进行预训练后,开源模型在HLE验证数据集的STEM子集上的性能较基线模型提升了15%,p值为0.045,表明该数据集具备用于模型训练的潜在价值。我们已向AI研究社区开源了部分数据集。

英文摘要

Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonomy distributions, (3) multiple choice question format that is misaligned with how scientists use AI in the real world, and (4) inaccurate answers and rationales partially led by a contest-based data collection and a time-bound review process. In this study, we present 'Expert-validated STEM QA', a high-quality, expert-validated STEM dataset (N=398) in Physics, Chemistry, Biology, and Mathematics, created by 241 domain experts. We (1) carefully designed a taxonomy with balanced distribution, (2) vetted question contributors with quality-driven incentive, (3) conducted multiple rounds of reviews with revisions validated by domain experts based on consensus, and (4) created the dataset in verifiable question and answer format. Our study demonstrated low performance ($<25\%$) of frontier AI models on the dataset as a benchmark. Post-training on a separate, private version of the dataset (N=2,000) increased performance of the open source model by $15\%$ relative to the baseline model (p=0.045) on the STEM subset of HLE-verified dataset, indicating potential utility of the dataset for model training. We have open-sourced a portion of our dataset for the AI research community.

CommentsWe have open-sourced a portion of our dataset for the AI research community at https://huggingface.co/datasets/TuringEnterprises/Open-RL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑