arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于SSMD的机器学习性能指标的高通量筛选实验命中选择

Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays

Xiaohua Douglas Zhang

arXiv 2608.07609首次发表:更新:

发表机构

University of Kentucky(肯塔基大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对高通量筛选的单次重复数据稀疏问题,提出基于SSMD的框架推导分类性能指标,验证其在丙型肝炎病毒siRNA筛选中可得到等效命中集,连接HTS统计与机器学习评估理论。

AI 中文摘要

高通量筛选(HTS)实验是早期药物发现的核心,但常受限于极端的数据稀疏性,因为初级筛选通常仅对每个测试物质使用单次重复。这种稀疏性使得传统机器学习性能指标,如灵敏度、特异度和受试者工作特征曲线下面积(AUROC),难以通过经验估计,因为它们需要足够规模的标记样本。在此,我们提出一种基于模型的框架,该框架从严格标准化均差(SSMD,一种成熟的HTS效应量参数)推导这些分类指标。在高斯等方差假设下,我们推导了将SSMD与Youden最优灵敏度和特异度、以及预设特异度下的灵敏度关联的闭式关系,从非中心t分布得出显式估计量和精确置信区间,即使在单次重复设计下也适用。与经典统计功效不同,经典统计功效无论组间真实非零差异多小,都会随样本量增大趋近于1;而SSMD推导的灵敏度会收敛到有限的总体值,反映两组间的真实分离程度,使其成为命中选择更有意义和稳定的性能度量。我们在包含约22000个单次重复测量的丙型肝炎病毒初级siRNA筛选中验证了该框架的实用性,结果显示SSMD、AUROC和基于灵敏度的阈值产生等效且可解释的命中集。本研究连接了经典HTS统计学与机器学习评估理论,为超低重复筛选工作流中的分类性能估计提供了统计上严谨、可重复的方法。

英文摘要

High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-learning performance metrics, such as sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC), difficult to estimate empirically because they require adequately sized labeled samples. Here, we introduce a model-based framework that derives these classification metrics from the strictly standardized mean difference (SSMD), a well-established HTS effect-size parameter. Under a Gaussian equal-variance assumption, we derive closed-form relationships linking SSMD to Youden-optimal sensitivity and specificity, and sensitivity at a preset specificity, yielding explicit estimators and exact confidence intervals from the noncentral t-distribution, even under single-replicate designs. Unlike classical statistical power, which approaches 1 as sample size grows regardless of how small the true non-zero difference between group means is, the SSMD-derived sensitivity converges to a finite population value that reflects the true degree of separation between two groups, making it a more meaningful and stable performance measure for hit selection. We demonstrate the utility of this framework in a hepatitis C virus primary siRNA screen comprising approximately 22,000 single-replicate measurements, showing that SSMD, AUROC, and sensitivity-based thresholds yield equivalent and interpretable hit sets. This work bridges classical HTS statistics and machine-learning evaluation theory, providing a statistically principled, reproducible way to estimate classification performance in ultra-low-replication screening workflows.

Comments20 pages, 2 figures

Journal refSLAS Discovery 44 (2026), 100340

DOI:10.1016/j.slasd.2026.100340

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑