arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自然语言引导的、与生成器无关的蛋白结合剂筛选

Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design

Gyubok Lee, Kiwoong Yoo, Jimin Seo, Jiyoun Kim, Kyunghoon Hur, Edward Choi

arXiv 2608.20755首次发表:更新:

发表机构

Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST); LG AI Research; Seoul National University; Korea Electronics Technology Institute (KETI)(韩国科学技术院金在哲人工智能研究生院; LG人工智能研究院; 首尔大学; 韩国电子技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出用LLM生成排序策略,从已生成的蛋白结合剂池中筛选候选,在10目标和3目标测试中均优于基线,为蛋白结合剂筛选提供了可解释的后处理方法。

AI 中文摘要

现代从头设计工作流会生成大量候选蛋白结合剂,但湿实验室验证能力有限,使得筛选成为主要瓶颈。本研究探讨大型语言模型(LLMs)是否能基于预计算的结构置信度和界面质量代理分数生成多指标排序策略。本文不提出新的蛋白结合剂设计流程,而是聚焦于生成后结合剂筛选:利用一组预计算的代理分数,从已生成的结合剂池中选出最终的前K个候选。在10个目标的留出拆分测试中,5个采样的全局迭代gpt-4o策略的平均表现达到0.589的Recall@10,较最强的单特征固定基线Protenix binder ipTM(Recall@10为0.571)有小幅提升。在包含尼帕病毒(Nipah)、RBX1和TREM2的3个目标留出子集中,目标条件迭代gpt-5.4策略达到LLM的最强表现,Recall@10为0.519,NDCG@10为0.583。这些结果表明,LLM生成的排序策略可作为可解释的生成后决策层,用于组合异构代理指标以从大型候选池中优先选择结合剂。

英文摘要

Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisting a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structural-confidence and interface-quality proxy scores. Rather than proposing a new protein binder design pipeline, we focus on post-generation binder shortlisting: selecting the final top-K candidates from already generated binder pools using a shared panel of precomputed proxy scores. On the 10-target held-out split, averaging performance over five sampled global iterative gpt-4o policies reaches 0.589 Recall@10, modestly improving over the strongest single-feature fixed baseline, Protenix binder ipTM, which reaches 0.571 Recall@10. On the 3-target held-out subset comprising Nipah, RBX1, and TREM2, target-conditioned iterative gpt-5.4 policies reach the strongest LLM performance, with 0.519 Recall@10 and 0.583 NDCG@10. These results suggest that LLM-generated ranking policies can act as an interpretable post-generation decision layer for combining heterogeneous proxy metrics to prioritize binders from large candidate pools.

CommentsAccepted at ICML 2026 Workshop on Generative and Agentic AI for Biology

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑