发表机构
National University of Singapore; University of Oxford(新加坡国立大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出全模态AI科学家OmniScientist,通过端到端全模态流水线处理多学科异构原始证据,在36个真实案例中完成从原始数据到手稿的全流程,性能优于仅用预计算特征的系统,为通用AI科学家提供可行方案。
AI 中文摘要
近期基础模型的进展已使AI科学家能够自动化越来越完整的研究工作流程,从假设生成、代码执行到论文手稿准备。然而仅工作流程覆盖不足以获取科学发现所需的全部证据,现有系统通常仅对文本、代码、标签或预计算摘要进行推理,使得智能体无法利用对科学发现具有决定性作用的空间、时间、跨通道及过程关系。我们提出OmniScientist,一种端到端的全模态AI科学家,可直接从异构原始证据开展多学科研究。感知层及3个分别负责构思、实验和撰写的自主智能体在确定性流水线中运行,允许观测结果在整个研究生命周期中塑造研究问题、实验决策和最终结论。该系统通过代码运行构思、严谨性和结论检查,强制实施新颖性筛选、统计有效性、执行溯源和数值可追溯性。我们在36个真实案例上评估OmniScientist,这些案例涵盖5个学科领域、4类科学证据,模态包括图像、信号、音频、视频、三维结构、轨迹、表格、公式和图表。该系统在全部36个案例中完成从原始数据到编译手稿的完整路径,采用参考推理主干时获得平均论文评分为6.3。与仅接收预计算标量特征的盲态变体配对比较,直接感知提升了全部7个评估维度,并在85%的成对判断中胜出。这些结果表明,全生命周期感知对基于证据的科学发现至关重要,为构建通用型AI科学家提供了可行路径。
英文摘要
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.
Comments30 pages, 13 figures, 19 tables. Project page: https://omni-scientist.github.io/