arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17271cs.AI

ASI-Bench:人工智能超级智能的黎明

ASI-Bench: At the Dawn of Artificial Superintelligence

Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao W… 展开作者

Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员推出首个联合评估AI创新探索与自主科学执行能力的基准ASI-Bench,逐步减少人类方法指导测试AI自主科研能力,发现当前AI仍严重依赖人类指导,该基准面向全球开放邀请贡献。

中文摘要 AI 辅助

人工智能超级智能(ASI)要求AI超越掌握现有知识,转向探索未知、创造新知识并将新想法转化为可验证的结果。然而,当前AI系统的能力仍主要建立在学习、压缩和应用现有人类知识的基础上。因此,现有的基准测试主要评估AI能否基于所学知识给出正确答案,或在大量人类指导下完成任务。为此,我们推出ASI-Bench,这是首个在通用研究领域联合评估AI系统创新探索与自主科学执行能力的基准,也是首个在同一研究项目中逐步减少人类方法指导,以测试AI能自主推进到何种程度的基准。ASI-Bench由40多位专家耗时31000多小时构建,包含11个科学领域的60个项目级研究任务,并逐步减少方法指导,以测试AI能否自主选择方法、开展研究并产出可验证结果。所有任务均经过专家评审、AI辅助审核、沙箱执行和评分者验证。在18种最先进的智能体-模型配置中,平均得分从提供完整方法指导时的50.91,降至仅指定方法时的29.10,再到智能体必须自行确定方法时的26.62。这一显著下降表明,当前系统仍严重依赖人类指导,距离自主开展端到端的项目级科学研究仍有很长的路要走。ASI-Bench面向全球开放,我们邀请各地研究人员和建设者贡献新任务、挑战当前AI的极限,助力加速人类迈向人工智能超级智能的集体进程,网址为this https URL。

英文摘要

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.

发表机构

  • Tsinghua University(清华大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • Harvard University(哈佛大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Michigan(密歇根大学)
  • University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Boston University(波士顿大学)
  • University of Queensland(昆士兰大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Flatiron Institute(弗拉蒂伦研究所)
  • Microsoft Research(微软研究院)
  • AG2 AI(AG2 AI公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑