TopU-LBVS:一个用于基于配体虚拟筛选的现实多目标基准
TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening
浏览论文内容
中文总结 AI 辅助
针对现有LBVS基准的缺陷,构建了覆盖93个靶标、含硬阴性诱饵的多目标基准TopU-LBVS,提供三种固定协议,并证明硬阴性筛选下基线性能显著下降。
中文摘要 AI 辅助
基于配体的虚拟筛选(LBVS)是早期药物发现中实用的初筛工具,但现有基准可能因随机阴性样本、简单诱饵、目标覆盖有限以及非标准化的评估协议而高估性能。我们提出了TopU-LBVS,一个在硬阴性筛选条件下用于LBVS的多目标基准。从精选的ChEMBL~35生物活性数据出发,TopU-LBVS覆盖了7个蛋白质类别中的93个蛋白质靶标,并构建了具有属性匹配、结构相似的诱饵、固定1:40活性物与诱饵比率的靶标特异性筛选库。库中包含约400至10,000个化合物,旨在减少简单的物理化学性质和最近邻指纹捷径。TopU-LBVS提供了三种固定协议。TopU-LBVS-full评估了ChEMBL$^\ast \rightarrow$ TopU在所有93个靶标上的泛化能力。TopU-LBVS-low评估了硬阴性分布内的低数据TopU $\rightarrow$ TopU学习。TopU-LBVS-mini提供了一个紧凑的七靶标协议,带有成对的随机诱饵对照,仅改变测试诱饵,从而能够进行低成本开发并直接测量随机ChEMBL$^\ast$与TopU诱饵之间的差距。在跨越指纹方法、分子GNN、指纹混合和现代分子模型的十个参考基线中,随机诱饵评估下的性能在硬阴性筛选中急剧下降。我们发布了数据、固定划分、评估代码和基线实现,以便对未来LBVS和分子表示学习方法进行可复现的比较。代码和数据可在以下网址获取:此https URL和此https URL。
英文摘要
Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL~35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts. TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL$^\ast \rightarrow$ TopU generalization across all 93 targets. TopU-LBVS-low evaluates low-data TopU $\rightarrow$ TopU learning within the hard-negative distribution. TopU-LBVS-mini provides a compact seven-target protocol with a paired random-decoy control that changes only the test decoys, enabling low-cost development and direct measurement of the gap between random ChEMBL$^\ast$ and TopU decoys. Across ten reference baselines spanning fingerprint methods, molecular GNNs, fingerprint hybrids, and modern molecular models, performance under random-decoy evaluation degrades sharply under hard-negative screening. We release data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of future LBVS and molecular representation learning methods. Code and data are available at https://github.com/topu-benchmark/topu-lbvs and https://huggingface.co/datasets/topu-benchmark/topu-lbvs.
发表机构
- UT Dallas(达拉斯德州大学)
- National Institute of Biological Sciences(国家生物科学研究所)
机构由 AI 辅助整理,请以论文原文为准。