arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30517cs.AIcs.CL

ScienceArena:在最新科学奥林匹克竞赛上对大语言模型(LLM)进行基准测试

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu, Yaoming Li, Linfeng Hao, Shuyang Hou, Zijian Guo, Xinrui Zhang, Yuntian Zhao, Zhengyang Wang, Wenrui Liu, Yu… 展开作者

Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu, Yaoming Li, Linfeng Hao, Shuyang Hou, Zijian Guo, Xinrui Zhang, Yuntian Zhao, Zhengyang Wang, Wenrui Liu, Yuhan Wu, Tong Yang, Lin Sun, Xiangzheng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出ScienceArena基准,涵盖多类科学竞赛,经专家审核构建并校准LLM评判者,评估14个LLM发现顶尖模型获奖牌级分数,化学与长程一致性仍为瓶颈。

中文摘要 AI 辅助

基准测试饱和与数据污染日益模糊前沿大语言模型(LLM)的真实科学推理能力。我们推出ScienceArena,这是一个源自13项公开科学竞赛的奥林匹克风格基准,涵盖物理、化学、生物领域,包括2025-2026年国际物理奥林匹克竞赛(IPhO)、2025-2026年国际化学奥林匹克竞赛(IChO)、2023年国际生物奥林匹克竞赛(IBO)、2026年美国物理奥林匹克竞赛(USAPhO)以及2025年美国国家化学奥林匹克竞赛(USNCO)。其开放式多步骤问题采用过程信用评分标准,使得准确评分颇具难度。我们构建ScienceArena时采用了经专家审核的数字化流程,将官方考试、图表、答案及评分标准转换为结构化条目,并由奥林匹克奖牌得主验证。为将评估规模扩展至成本高昂的人工评分之外,我们针对IPhO和IChO中5个模型的存档答案,将“LLM作为评判者”与奖牌得主的真实结果进行校准;两个表现出色的评判者与专家总分的差距在1分以内。奖牌得主的标注显示,失败通常源于视觉定位、结构保真度及全局问题把控,而非术语缺失。通过交错求解评估14个近期LLM,我们发现顶尖模型在若干公开国际竞赛中获得了与奖牌相当的评分标准分数,而化学领域及长程一致性仍是关键瓶颈。我们提供交互式演示。

英文摘要

Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.

发表机构

  • Qiyuan Tech(奇安信科技)
  • Tsinghua University(清华大学)
  • The University of Hong Kong(香港大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑