发表机构
University of California, Los Angeles; Harvard University; Wyss Institute for Biologically Inspired Engineering; Columbia University; Terasaki Institute for Biomedical Innovation; Georgia Institute of Technology; The University of Tokyo; University of Southern California; University of Michigan; National University of Singapore; Massachusetts Institute of Technology; ETH Zurich; University of Chicago; University of Virginia; Northwestern University; Duke University; Institute for Basic Science; Yonsei University; Boston University(加州大学洛杉矶分校; 哈佛大学; 怀斯生物启发工程研究所; 哥伦比亚大学; 寺崎生物医学创新研究所; 佐治亚理工学院; 东京大学; 南加州大学; 密歇根大学; 新加坡国立大学; 麻省理工学院; 苏黎世联邦理工学院; 芝加哥大学; 弗吉尼亚大学; 西北大学; 杜克大学; 基础科学研究院; 延世大学; 波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BioEVAL是一个全球多机构博士级基准,涵盖608项生物工程评估,包括多项选择、文献综合和多模态任务,评测多种LLM,最高准确率达90%,旨在衡量实验推理能力。
AI 中文摘要
大型语言模型(LLM)在通用推理方面已展现出历史性突破,并在生物医学科学领域取得了早期成功。然而,现有的LLM基准测试侧重于事实性回忆,对模型在前沿和多模态任务上的性能提供的洞察有限。我们组建了BioEVAL(生物工程AI与LLM验证),这是一项全球性的多机构倡议,旨在评估生物工程(BE)各子领域的实验推理能力。BioEVAL涵盖11个主要的BE子领域及一组未分类项目,汇集了22个研究团队,创建了一个博士级基准,包含608个评估项目:1)380道多项选择题(MCQs,审计后保留359道),2)218项文献综合任务,以及3)10个涉及实验图像解释的多模态问题。基准项目在评估前经过了编写团队的专家审查和集中质量控制。评估后,对准确率最高和最低的MCQ项目进行的盲法跨团队共识审计标记了21个问题需要修订或移除;这些问题被扣留,所有报告的MCQ结果均基于保留的359个项目计算。我们评估了多种云规模的基础/多模态模型(如ChatGPT、Gemini和Grok)以及适合在消费级GPU上进行推理的本地可部署模型。模型在MCQ上达到了高达90%的最高准确率,在文献综合上的相似度得分为0.72,在小样本多模态推理问题上的准确率为80%,且各子领域间的性能差异显著。排行榜排名表征了所评估的BE任务类别中当前的能力、局限性和发展优先级。BioEVAL作为一个可扩展的基准得以维护,并采用标准化协议,以持续进行专家项目贡献和模型评估。
英文摘要
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major BE subfields plus a set of uncategorized items, bringing together 22 research groups to create a PhD-level benchmark comprising 608 evaluation items: 1) 380 multiple-choice questions (MCQs, 359 retained after audit), 2) 218 literature synthesis tasks, and 3) 10 multimodal problems with experimental image interpretation. Benchmark items underwent authoring-group expert review and centralized quality control before evaluation. Following evaluation, a blinded cross-group consensus audit of the highest- and lowest-accuracy MCQ items flagged 21 questions for revision or removal; these were withheld, and all reported MCQ results are computed on the 359 retained items. We evaluated diverse cloud-scale foundation/multimodal models (e.g., ChatGPT, Gemini, and Grok) and locally deployable models suitable for inference on consumer-grade GPUs. Models achieved the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with substantial performance variation across subfields. Leaderboard rankings characterize current capabilities, limitations, and development priorities across the evaluated BE task categories. BioEVAL is maintained as an extensible benchmark with standardized protocols for continuing expert item contribution and model evaluation.