发表机构
Google DeepMind; Duke University; Columbia University; Google Research; Texas A&M University(谷歌DeepMind; 杜克大学; 哥伦比亚大学; 谷歌研究; 德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究扩展并验证了基于Gemini的多智能体系统Co-Scientist,其可在材料、生物、计算机科学领域推进闭环科学工作流,提升研究效率并减少幻觉抄袭。
AI 中文摘要
我们提出了Co-Scientist的扩展方案及全面的现实世界验证,Co-Scientist是一种基于Gemini的多智能体系统,旨在加速从假设生成、实验到手稿生成的端到端科学研究。该专用配置超越了计算机模拟假设生成,将Co-Scientist转变为基于执行的研究伙伴,推进材料科学、生物学和计算机科学领域的闭环科学工作流。在材料科学领域,Co-Scientist与半自动化化学气相沉积反应器对接,设计了MXenes的安全前驱体路线;实验执行产出了层状二维材料,其与Ti3C2Tx MXene晶格具有关键结构相似性,尽管需进一步实验确认原子结构。借助Gemini 3 Deep Think实现快速的实验室闭环执行,它还在数分钟内针对实验室约束定制了生长配方,实现了单层MoS2、MoSe2和WS2半导体的单次生长。在生物学领域,Co-Scientist从稀疏成像数据中预测了工程化大肠杆菌在诱导剂(IPTG)梯度下的涌现 swarm 表型,定量匹配了未发表的湿实验室形态学测量结果。在计算机科学领域,Co-Scientist自主发现了一种推理时缩放架构,在HealthBench(硬核与专业子集)上的表现优于6个前沿模型,同时在盲法医师评估下降低了潜在临床伤害。最后,针对端到端生成论文的双盲研究,由30名领域专家完成450次评审,结果表明Co-Scientist的可靠性模块减少了幻觉和抄袭,同时提升了研究安全性。综上,这些结果证明了在能够加速现实世界科学发现的闭环多智能体科学AI系统方面取得了进展。
英文摘要
We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypothesis generation, this specialized configuration transitions Co-Scientist into an execution-grounded research partner advancing closed-loop scientific workflows across materials science, biology, and computer science. In materials science, Co-Scientist interfaced with a semi-automated chemical vapor deposition reactor to design a safe precursor route for MXenes; experimental execution produced a lamellar 2D material sharing key structural similarities with the Ti3C2Tx MXene lattice, although further experiments are needed to confirm the atomic structure. Leveraging Gemini 3 Deep Think for rapid, lab-in-the-loop execution, it also tailored growth recipes to laboratory constraints in minutes, enabling single-attempt growth of monolayer MoS2, MoSe2, and WS2 semiconductors. In biology, Co-Scientist predicted emergent swarming phenotypes of engineered E. coli across inducer (IPTG) gradients from sparse imaging data, quantitatively matching unpublished wet-lab morphological measurements. In computer science, Co-Scientist autonomously discovered an inference-time scaling architecture that outperformed six frontier models on HealthBench (Hard and Professional) while reducing potential clinical harm under blinded physician evaluation. Finally, a double-blind study of end-to-end generated papers with 30 domain experts across 450 reviews demonstrates that Co-Scientist's reliability modules reduce hallucination and plagiarism while improving research safety. Together, these results demonstrate progress toward closed-loop multi-agent scientific AI systems capable of accelerating real-world scientific discovery.