发表机构
University of California San Diego; Texas A&M University; Carnegie Mellon University; Mohamed bin Zayed University of Artificial Intelligence(加州大学圣迭戈分校; 德克萨斯A&M大学; 卡内基梅隆大学; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HypoEvolve利用代际遗传算法协调多个大语言模型智能体协作,通过明确种群更新规则,在药物重定位的34种癌症类型评估中超越六种基线,提升假说质量。
AI 中文摘要
科学智能体通过综合证据、评估提案和开发新解释来促进假说发现。近期系统将科学智能体与进化搜索相结合,通过批判、比较和修订来实现。然而,不同形式的智能体协作如何影响假说质量仍是一个悬而未决的问题。回答这一问题需要将智能体的科学能力与其协作的影响分离开来。因此,一个框架必须保留智能体的科学角色,并支持组合、修订和保留假说的规则。基于这一观点,我们提出了HypoEvolve,它通过对假说种群的连续更新使协作变得明确。具体而言,我们提出了一种代际遗传算法来协调专门的大语言模型(LLM)智能体,这些智能体整合机制论证、重新考虑假设并评估证据和可检验性。每一代都规定了科学判断和新提案如何重塑种群,从而使协作对假说质量的影响可直接检验。此外,我们围绕具有科学意义的假说设计评估,这些假说解释了一个拟议干预措施如何发挥作用。药物重定位将这些解释与针对外部证据评估的靶点级生物学声明联系起来。具体而言,我们将DepMap和Open Targets调整为互补的外部度量,这些度量基于实验、遗传和临床证据。在34种癌症类型中,HypoEvolve在两项度量上均取得了六种基线中最高的分数。DepMap选择性达到0.171,而最强基线为0.115。相对于单次生成的优势也泛化到保留的癌症类型。HypoEvolve推进了自主科学的愿景,其中AI研究团队实现了超越单个模型的发现能力。
英文摘要
Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses. Building on this view, we introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, we propose a generational genetic algorithm to coordinate specialized large language model (LLM) agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making collaboration effects on hypothesis quality directly testable. Moreover, we design our evaluation around scientifically meaningful hypotheses that explain how a proposed intervention could work. Drug repurposing links these explanations to target-level biological claims assessed against external evidence. Specifically, we adapt DepMap and Open Targets into complementary external measures grounded in experimental, genetic, and clinical evidence. Across 34 cancer types, HypoEvolve achieves the highest scores against six baselines on both measures. DepMap selectivity reaches 0.171, versus 0.115 for the strongest baseline. Gains over single-pass generation also generalize to held-out cancer types. HypoEvolve advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.
Comments22 pages, 8 figures, 5 tables