发表机构
University of Oxford; University College London; Technical University of Munich; Oxford-Suzhou Centre for Advanced Research, China(牛津大学; 伦敦大学学院; 慕尼黑工业大学; 牛津-苏州先进研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LabFactory框架让AI构建者将科学简报转化为可执行AI实验室,通过独立主机评估交付系统,在28个构建、33个子测试中全部超过参考值,证明AI可端到端实现科学任务。
AI 中文摘要
科学任务指定了所需的能力,但实现该能力通常需要构建一个针对任务定制的计算系统——包括获取数据、设计表示、训练模型、实现工具,以及决定在推理时如何使用它们。我们提出了LabFactory,这是一个框架,其中AI构建者将科学简报转化为可执行的AI实验室:一个任务特定的求解器,将模型、知识资源、工具和控制器集成在固定接口之后。构建者在计量工作空间中开发并打包实验室;随后,一个独立的主机在保留的输入上执行交付的工件,参考标签保持在求解器输入接口之外,并根据任务的协议对其输出进行评分。这使得交付的系统(而非构建者对其进展的描述)成为评估的对象。我们记录了七个科学任务类别中的28个精选构建——从分子和基因组预测到生理信号、临床决策支持和生物医学文本——其交付的实验室在主机端执行下,在所有33个子测试中均超过了其配置的参考值。其中十个包含在构建期间拟合的预测模型;其余则围绕固定的平台LLM组装检索系统、可执行分析环境和工具驱动的工作流。它们共同表明,AI代理可以将科学简报一直带到可工作的实验室,并且在构建结束后,该实验室仍可被调用、检查和验证。
英文摘要
Scientific tasks specify a desired capability, but realizing it often requires building a computational system tailored to the task---acquiring data, designing representations, training models, implementing tools, and deciding how they are used at inference. We present, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface. The builder develops and packages the lab in a metered workspace; a separate host then executes the delivered artifact on held-out inputs, with reference labels kept outside the solver's input interface, and scores its outputs under the task's protocol. This makes the delivered system, rather than the builder's account of its progress, the object of evaluation. We document 10 selected constructions across six scientific task categories---from molecular and genomic prediction to medical imaging, clinical decision support, and biomedical text---whose delivered labs exceeded their configured reference values on all 12 subtests under host-side execution. Four contain predictive models fitted during construction; the others assemble executable analysis environments, knowledge resources, and tool-driven workflows around a fixed platform LLM. Together they show that an AI agent can carry a scientific brief all the way to a working lab that can still be invoked, inspected, and checked after construction ends.