IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
IMProofBench:在研究级数学证明生成上对人工智能进行基准测试
Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck, Jeremy Feusi, Tim Gehrunger, Raphael Appenzeller, Pieter Belmans, Alessio Bottini, Jim Bryan, João Camarneiro, Ana Cannas da Silva, Niklas Canova, Ana-Maria Castravet, Timo de Wolff, Claudio Fontanari, Filippo Gaia, Baran Hashemi, Daniel Holmes, David Holmes, Aitor Iribar Lopez, Victor Jaeck, Martina Jørgensen, Steven Kelk, Martijn Kool, Stefan Kuhlmann, Adam Kurpisz, Johannes Lengler, Chiara Meroni, Ingmar Metzler, Martin Möller, Samuel Muñoz-Echániz, David Muñoz-Lahoz, Robert Nowak, Georg Oberdieck, Daniel Platt, Dylan Possamaï, Gabriel Ribeiro, Aluna Rizzoli, Daria Sakhanda, Raúl Sánchez Galán, Zheming Sun, Diaaeldin Taha, Josef Teichmann, Richard P. Thomas, Henk van der Pol, Michel van Garrel, Charles Vial, Ignacio Barros, Benjamin Doerr, Peter Grünwald, Henry Liu, David Martins, Aleksandar Mijatović, Sergej Monavari, Marc Roth, Patrick Schnider, Yannik Schuler, Pim Spelier, Yuuji Tanaka, Ronald van Luijk
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
Aarhus University(奥胡斯大学)
Commentsv2: benchmark expanded from 39 to 77 problems; evaluation extended to 14 models including GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6; new analyses (IRT-based score aggregation, inter-rater reliability, tool/token usage, non-agentic ablation); contributor author list updated
机构
*
School of Computer and Information Engineering, Henan University(河南大学计算机与信息工程学院)
;
Department of Statistics, LMU Munich(慕尼黑大学统计系)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
机构
*
State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, CAS(中国科学院数学与系统科学研究院数学科学国家重点实验室)
;
School of Mathematical Sciences, University of Chinese Academy of Sciences(中国科学院大学数学科学学院)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学高等研究院)
RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning
RolloutPipe: 分离式在线策略LLM强化学习中的流水线化Rollout与训练重叠
Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu
机构
*
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Southern University of Science and Technology(南方科技大学)
AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database
AutoMathKG:基于LLM和向量数据库的自动化数学知识图谱
Rong Bian, Yu Geng, Zijian Yang, Bing Cheng
机构
*
Academy of Mathematics and Systems Science(数学与系统科学研究院)
;
Chinese Academy of Sciences(中国科学院)
;
School of Mathematical Sciences(数学科学学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
AMSS Center for Forecasting Science(数学与系统科学研究院预测科学中心)
;
State Key Laboratory of Mathematical Science(数学科学国家重点实验室)
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
AfriqueLLM: 数据混合与模型架构如何影响非洲语言的持续预训练
Hao Yu, Tianyi Xu, Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, David Ifeoluwa Adelani
机构
*
McGill University(麦吉尔大学)
;
Mila-Quebec AI Institute(魁北克AI研究所)
;
LMU Munich & Munich Center for Machine Learning(慕尼黑大学及慕尼黑机器学习中心)
;
Microsoft AI for Good Research Lab(微软AI for Good研究实验室)
Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery
Science Earth: 迈向面向AI原生科学发现的行星级操作系统
Zhe Zhao, Haibin Wen, Yingcheng Wu, Jiaming Ma, Yifan Wen, Jinglin Jian, Jiacheng Ge, Xiangru Tang, Bo An, Ming Yin, Sanfeng Wu, Mengdi Wang, Le Cong
机构
*
Department of Pathology, Department of Genetics, Stanford University School of Medicine(病理学系、遗传学系,斯坦福大学医学院)
;
Princeton AI Lab, Department of Electrical & Computer Engineering, Princeton University(普林斯顿人工智能实验室、电气与计算机工程系,普林斯顿大学)
;
Scripps Research, La Jolla, CA, USA(斯克里普斯研究机构,洛杉矶,加利福尼亚州,美国)
;
Division of Biostatistics, Department of Population Health, New York University Grossman School of Medicine(生物统计学部、人口健康系,纽约大学格罗斯曼医学院)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
;
Department of Computer Science, Yale University(计算机科学系,耶鲁大学)
;
Department of Physics, Princeton University(物理系,普林斯顿大学)
CommentsWithdrawn by the authors. (1) The author list and authorship roles had not been finalized and agreed upon by all listed authors prior to submission. (2) The specific contribution of the system in the K3 synchronization example (Section on Kuramoto/nonlinear physics) requires further validation before it can be reported. The authors are addressing both points and may resubmit a corrected version.
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Jinesis Lab, University of Toronto & Vector Institute(Jinesis实验室,多伦多大学及向量研究所)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Princeton University(普林斯顿大学)
;
Cornell University(康奈尔大学)
;
The University of Tokyo(东京大学)
;
RIKEN AIP(日本理化学研究所AIP)
;
Max Planck Institute for Intelligent Systems, Tübingen, Germany(德国图宾根最大计划智能系统研究所)
;
EuroSafeAI