arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11877q-bio.QMcs.AIcs.CLq-bio.GN

生物学在环:CRISPR筛选中的摊销自适应命中发现

Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

Carl Edwards, Edward De Brouwer, Xiner Li, Namkyeong Lee, Ehsan Hajiramezanali, Anne Biton, Sara Mostafavi, Gabriele Scalia

首次发表
浏览论文内容

中文总结 AI 辅助

针对CRISPR筛选中的自适应命中发现,提出AssayBench-Loop基准和AssayLoop框架,结合Transformer策略与LLM先验,实现5.67倍富集和27.7%命中恢复。

中文摘要 AI 辅助

许多生物学发现问题需要在受限预算下顺序选择实验。CRISPR筛选是一个突出的例子,因为穷举性扰动测试通常不可行,候选扰动必须在多轮实验中按优先级排序。尽管这一问题十分重要,现有的自适应命中发现基准在规模和多样性方面仍然有限。在此,我们引入了AssayBench-Loop,一个大规模的自适应命中发现基准,包含跨越五个表型类别的1,389个CRISPR筛选。除了支持系统性评估外,其规模使得我们能够跨历史实验学习获取策略。基于这一资源,我们提出了AssayLoop,一个序贯实验设计框架,结合了AssayFormer(一种基于Transformer的摊销获取策略,在历史筛选中训练以从实验反馈中适应)以及通过自适应交接从LLM衍生的生物学先验。在这一视角下,已完成的实验成为训练数据,用于学习累积证据应如何指导下一步测试什么,而LLM提供先验生物学知识以启动搜索。我们进一步引入了AssayLLM,表明同一原则可以通过任务特定的后训练直接扩展到LLM。在时间上留出的筛选中,AssayLoop实现了相对于随机选择的5.67倍富集,并在检测约5%的候选文库后恢复了27.7%的命中,优于现有的自适应设计方法和独立的LLM,以及仅使用AssayFormer的方法。性能随着历史训练数据的增加而提高,并迁移到训练中排除的表型类别。这些结果证明了跨历史实验学习获取策略并将其与广泛的生物学先验相结合对于高效自适应命中发现的价值。

英文摘要

Many biological discovery problems require experiments to be selected sequentially under constrained budgets. CRISPR screening is a prominent example, as exhaustive perturbation testing is often infeasible and candidate perturbations must instead be prioritized over multiple experimental rounds. Despite the importance of this problem, existing benchmarks for adaptive hit discovery remain limited in scale and diversity. Here, we introduce AssayBench-Loop, a large-scale benchmark for adaptive hit discovery comprising 1,389 CRISPR screens across five phenotype categories. Beyond enabling systematic evaluation, its scale makes it possible to learn acquisition strategies across historical experiments. Building on this resource, we introduce AssayLoop, a sequential experimental design framework combining AssayFormer, a transformer-based amortized acquisition policy trained across historical screens to adapt from experimental feedback, with LLM-derived biological priors through an adaptive handoff. In this view, completed experiments become training data for learning how accumulated evidence should guide what to test next, while LLMs provide prior biological knowledge to seed the search. We further introduce AssayLLM, showing that the same principle can be extended directly to an LLM through task-specific post-training. On temporally held-out screens, AssayLoop achieves a 5.67-fold enrichment over random selection and recovers 27.7% of hits after assaying approximately 5% of the candidate library, outperforming existing adaptive-design methods and standalone LLMs, and AssayFormer alone. Performance improves with increasing historical training data and transfers to phenotype categories excluded from training. These results demonstrate the value of learning acquisition policies across historical experiments and combining them with broad biological priors for efficient adaptive hit discovery.

发表机构

  • Genentech(基因泰克)

机构由 AI 辅助整理,请以论文原文为准。

↑