arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01827cs.LG

验证拥堵下的科学发现:基于多保真度成对排序的方法

Scientific Discovery under Validation Congestion via Multi-Fidelity Pairwise Rankings

Kevin Tirta Wijaya, Alston Lo, Michael Sun, Wojciech Matusik, Vahid Babaei

首次发表
浏览论文内容

中文总结 AI 辅助

针对候选物多而实验验证能力有限的验证拥堵问题,提出PRISMS框架,利用多保真度专家成对排序进行科学设计筛选,无需大量实验数据,显著提升筛选效率和优化性能。

中文摘要 AI 辅助

现代计算方法现在能够以前所未有的规模提出候选分子、材料和其他科学设计,这造成了验证拥堵:候选物数量丰富,但用于物理评估它们的实验能力仍然稀缺。因此,发现新的科学设计越来越依赖于筛选:选择一小部分有前景的设计用于缓慢且昂贵的实验。现有的筛选方法通常依赖于预测绝对分数的数据驱动回归模型,但训练这些模型首先需要大量的实验数据。然而,有用的筛选信号不必以绝对测量的形式存在,因为科学设计的发现本质上是比较性的。在这里,我们提出筛选可以主要由专家成对排序驱动,这种排序更容易收集。专业知识可以来自计算工具或人类输入,具有多种保真度级别,从经验法则到智能体工作流和经验丰富的科学家。我们引入了PRISMS框架,该框架利用来自一个或多个专家的成对排序,可能跨越多个专业水平,来识别最有前景的候选物,而不依赖于数据饥渴的回归器。当专家在保真度和成本上有所不同时,PRISMS基于Fisher信息准则,将成对查询从低保真度排序器升级到高保真度排序器。在从固定药物发现库中选择设计的迭代筛选中,PRISMS在比仅回归的主动学习少约42%的轮次内实现了50%的前10名发现召回率,并且比没有选择性升级的基于排序的方法少约15%的轮次。在生成新设计而不限于预定义库的优化中,PRISMS实现的超体积比贝叶斯优化基线高约18.8%。

英文摘要

Modern computational methods can now propose candidate molecules, materials, and other scientific designs at an unprecedented scale, creating a validation congestion where candidates are abundant, but experimental capacity to physically evaluate them remains scarce. Discovering novel scientific designs has therefore become increasingly dependent on curation: selecting a small set of promising designs for slow and costly experiments. Existing curation methods typically rely on data-driven regression models that predict absolute scores, but training these models requires substantial experimental data to begin with. Yet, useful curation signals do not have to take the form of absolute measurements, as scientific design discovery is often comparative in nature. Here, we propose that curation can instead be primarily driven by expert pairwise rankings, which are substantially easier to gather. The expertise can come from computational tools or human input of multiple levels of fidelity, ranging from empirical rules of thumb to agentic workflows and experienced scientists. We introduce PRISMS, a framework that uses pairwise rankings from one or more experts, potentially spanning multiple levels of expertise, to identify the most promising candidates without relying on data-hungry regressors. When experts differ in fidelity and cost, PRISMS escalates pairwise queries from lower- to higher-fidelity rankers based on a Fisher-information criterion. In iterative screening that selects designs from fixed drug discovery libraries, PRISMS achieves 50% top-10 discovery recall in ~42% fewer rounds than regression-only active learning, and in ~15% fewer rounds than the ranking-based method with no selective escalation. In optimization that generates new designs without restriction to a predefined library, PRISMS achieves ~18.8% higher hypervolume than the Bayesian optimization baseline.

发表机构

  • University of Bonn(波恩大学)
  • Fraunhofer SCAI(弗劳恩霍夫算法与科学计算研究所)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑