发表机构
Friedrich-Alexander-Universität Erlangen-Nürnberg; Technical University of Munich; IEO European Institute of Oncology IRCCS(埃尔朗根-纽伦堡弗里德里希-亚历山大大学; 慕尼黑工业大学; IEO欧洲肿瘤研究所IRCCS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统审计PPI数据集中的多种偏差,提出结合相似性感知划分与偏差最小化负采样的Nextflow流水线,以缓解模型学习捷径而非生物信号的问题。
AI 中文摘要
蛋白质-蛋白质相互作用(PPI)数据库并不能忠实地反映生物学现实。相反,它们受到研究和技术偏差的影响,这些偏差扭曲了某些蛋白质和相互作用的属性。如果负数据集构建不当,机器学习模型可能会利用这些偏差作为学习捷径。到目前为止,在PPI数据集构建过程中引入的捷径仅被孤立地研究。在这里,我们系统地描述了PPI数据集中已报道的以及据我们所知先前未报道的偏差,这些偏差导致机器学习模型学习捷径而非生物信号。我们分析了专门的PPI数据库HIPPIE、IntAct和STRING,以及两个源自蛋白质数据银行(PDB)中三维结构信息的数据集。我们表明,随机数据划分引入了强烈的拓扑捷径。当去除训练-测试蛋白质重叠时,所得数据集仍然保留来自自身相互作用、分类学同一性和功能相关性的可用捷径,有趣的是,这些捷径的普遍性取决于数据来源。我们进一步表明,从一组高置信度非相互作用者中采样负样本(一个直观上吸引人的选择)会放大由功能相关性引起的捷径。为了检测和缓解这些偏差,我们提供了一个开放的Nextflow流水线,该流水线将相似性感知、数据损失最小化的数据集划分与偏差最小化的负采样相结合,两者均被表述为整数线性规划。其通过基于优化的负采样来量化偏差以最小化偏差的关键概念,原则上可以扩展到任何负候选池远大于正样本的机器学习问题,因此其意义也超出了PPI预测的范畴。
英文摘要
Protein-protein interaction (PPI) databases do not faithfully reflect biological realities. Instead, they are influenced by study and technical biases that distort certain protein and interaction attributes. Machine learning models can exploit these as learning shortcuts if the negative dataset is not constructed with care. So far, the shortcuts introduced during PPI dataset construction have only been examined in isolation. Here, we systematically characterize both reported and, to our knowledge, previously unreported biases in PPI datasets that lead machine learning models to learn shortcuts instead of biological signal. We analyze HIPPIE, IntAct, and STRING, dedicated PPI databases, as well as two datasets derived from 3D-structural information in the Protein Data Bank (PDB). We show that random data splitting introduces strong topological shortcuts. When train-test protein overlap is removed, the resulting datasets still retain usable shortcuts stemming from self-interactions, taxonomic identity, and functional relatedness, whose prevalence interestingly depends on the data source. We further show that sampling negatives from a set of high-confidence non-interactors, an intuitively appealing choice, can amplify the shortcut stemming from functional relatedness. To detect and mitigate these biases, we provide an open Nextflow pipeline that combines similarity-aware, data-loss-minimizing dataset splitting with bias-minimizing negative sampling, both formulated as integer linear programs. Its key concept of quantifying biases to minimize them through optimization-based negative sampling can, in principle, be extended to any machine learning problem where the pool of negative candidates is much larger than the positives and is thus of interest also beyond PPI prediction.