arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FrontierChallenge:科学工作流完成度评估

FrontierChallenge: Evaluating Scientific Workflow Completion

Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Brian Wang, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang

arXiv 2608.24979首次发表:更新:

发表机构

Apodex Team(Apodex团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出跨领域科学工作流基准FrontierChallenge,评估12个前沿模型的97项任务,发现高部分分数或完成声称无法可靠表明任务完成,需评估端到端执行与成果完整性。

AI 中文摘要

智能体越来越多地用于分析数据、执行代码并生成研究成果,但大多数基准测试侧重于最终答案、孤立程序或单一领域。本文介绍了FrontierChallenge,这是一个包含300个端到端科学工作流的跨领域基准。本文发布并评估了其中97个任务,涵盖量子化学、分子动力学、材料表征、分析化学、生命科学以及电化学/环境领域。每个任务提供固定输入并指定一组所需的科学成果。我们使用三种智能体框架评估了12个前沿模型。通过率(Pass Rate)衡量满足完整完成标准的任务比例,而平均分数(Avg. Score)捕捉部分进展。表现最佳的配置仅完成了97个发布任务中的20个,通过率为20.6%。部分进展转化为完整交付的情况在分析化学和电化学/环境领域尤为不佳:平均分数分别达到87.6和94.9,但最高通过率仅为4%和0%。在未通过的Claude Code轨迹中,75.5%仍以声称完成的语言结束。这些发现表明,高部分分数或自信的完成声称都不能可靠地表明科学任务已完全完成,凸显了需要同时评估端到端工作流执行和科学成果的完整性。

英文摘要

Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. Complementary HDS6 process scores correlate strongly with task outcomes, supporting FrontierChallenge as a benchmark of Heavy Duty Solver capabilities. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.

CommentsProject Website: https://apodexai.github.io/FrontierAgent/benchmarks/FrontierChallenge/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑