arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10595q-bio.BMcs.AI

DegradeQuery:面向上下文感知PROTAC降解预测的反事实元组预训练

DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction

  • School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
  • EasternDawn(东方黎明)
  • School of Computer Science, University of Nottingham Ningbo(诺丁汉大学宁波校区计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

Dong Xu, Zhangfan Yang, Jiantao Wu, Zexuan Zhu, Jianqiang Li, Junkai Ji

AI总结:

DegradeQuery是一种上下文感知PROTAC降解预测框架,通过反事实元组预训练利用标签缺失的PROTAC记录,在PROTAC-8K基准上AUC达0.9065、准确率0.8500,性能优于对比方法。

AI中文摘要:

蛋白水解靶向嵌合体(PROTACs)通过将靶蛋白募集至E3泛素连接酶来诱导蛋白降解,这使得降解是降解剂分子与其生物上下文的共同结果。尽管公共数据库包含数千条结构化的分子-靶标-E3记录,但仅对其中一小部分有降解测量值。因此,现有的监督方法未利用大多数已记录的化学生物关系。我们引入DegradeQuery,一种上下文感知预测框架,将这些标签缺失的记录转化为预训练信号。其反事实元组预训练目标是将已记录的元组与通过替换靶标、E3连接酶或两者形成的替代元组进行对比,使模型能够学习上下文关联而无需分配活性伪标签。随后对得到的表示进行微调,以从完整的分子-靶标-E3上下文预测降解。在官方PROTAC-8K基准上,DegradeQuery的受试者工作特征曲线下面积达0.9065,准确率为0.8500,优于对比方法。控制分析进一步表明,该改进主要归因于元组级预训练,仅使用标签缺失的记录即可实现,且与蛋白质语言模型表示互补。这些发现表明,标签不完整的PROTAC数据库包含有用的关系监督信号,为从稀缺实验标签中学习上下文感知的降解预测器提供了实用途径。

英文摘要:

Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. Although public databases contain thousands of structured molecule-target-E3 records, degradation measurements are available for only a small fraction of them. Existing supervised approaches therefore leave most recorded chemical-biological relationships unused. We introduce DegradeQuery, a context-aware prediction framework that converts these label-missing records into a pretraining signal. Its counterfactual tuple pretraining objective contrasts recorded tuples with alternatives formed by replacing the target, the E3 ligase, or both, enabling the model to learn contextual associations without assigning activity pseudo-labels. The resulting representation is then fine-tuned to predict degradation from the complete molecule-target-E3 context. On the official PROTAC-8K benchmark, DegradeQuery achieves an area under the receiver operating characteristic curve of 0.9065 and an accuracy of 0.8500, outperforming the compared methods. Controlled analyses further show that the improvement is primarily attributable to tuple-level pretraining, can be recovered using only label-missing records, and remains complementary to protein language model representations. These findings demonstrate that incompletely labeled PROTAC databases contain useful relational supervision and provide a practical route for learning context-aware degradation predictors from scarce experimental labels.

补充信息

↑