发表机构
Hunan University; Yuelushan Laboratory; Nanjing University of Posts and Telecommunications; Nanyang Technological University(湖南大学; 岳麓山实验室; 南京邮电大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ProMeta提出基于原型的图神经网络元学习框架,解决PROTAC降解活性预测中跨E3连接酶的少样本泛化问题,在CRBN到VHL基准上显著优于监督基线。
AI 中文摘要
蛋白水解靶向嵌合体(PROTACs)已成为一种变革性的治疗策略,通过泛素-蛋白酶体系统选择性降解历史上“不可成药”的靶点。尽管开发PROTAC降解活性计算预测器的努力日益增多,现有的监督方法仍严重受限于跨E3连接酶的数据稀缺和不平衡问题,限制了其在研究充分的连接酶情境之外的泛化能力。实践中,标记数据高度集中于少数连接酶(如CRBN和VHL),而大多数E3连接酶仍未被充分探索,但这对扩展靶向降解剂的设计空间至关重要。因此,开发能够在最少标记数据下实现稳健跨连接酶泛化的方法,对于提高计算PROTAC发现的实用价值至关重要。我们将跨E3连接酶的PROTAC降解活性预测重新构建为少样本元学习问题,并提出ProMeta,一种基于原型的图神经网络,通过源E3任务上的情景元学习进行训练,并通过支持条件推理在保留的目标E3任务上评估。ProMeta在不更新编码器的情况下进行推理,通过从最少的目标连接酶支持样本中动态估计类别原型。在CRBN到VHL基准上,ProMeta在K=2、Q=3下取得AUROC值0.796,在K=2、Q=5下取得0.883,分别比相应的监督GNN基线提高19.9%和6.8%。相同协议下反向VHL到CRBN迁移取得AUROC值0.702(K=2、Q=3)和0.821(K=2、Q=5),确认了双向适用性,同时揭示了方向和数据机制依赖性。综合这些结果,支持ProMeta作为在评估的支持/查询协议下进行跨连接酶少样本预测的实用框架。
英文摘要
Proteolysis-targeting chimeras (PROTACs) have emerged as a transformative therapeutic strategy that selectively degrades historically ''undruggable'' targets via the ubiquitin-proteasome system. Despite growing efforts to develop computational predictors of PROTAC degradation activity, existing supervised approaches remain severely challenged by data scarcity and imbalance across E3 ligases, limiting their ability to generalize beyond well-studied ligase contexts. In practice, labeled data are heavily concentrated on a few ligases (e.g., CRBN and VHL), while the majority of E3 ligases remain underexplored yet are critical for expanding the design space of targeted degraders. Developing methods that enable robust cross-ligase generalization with minimal labeled data is therefore essential for improving the practical utility of computational PROTAC discovery. We reformulate PROTAC degradation activity prediction across E3 ligases as a few-shot meta-learning problem and present ProMeta, a prototype-based graph neural network trained through episodic meta-learning on source-E3 tasks and evaluated on held-out target-E3 tasks through support-conditioned inference. ProMeta performs inference without updating the encoder by dynamically estimating class prototypes from minimal target-ligase support samples. On the CRBN-to-VHL benchmark, ProMeta achieves AUROC values of 0.796 under K=2, Q=3 and 0.883 under K=2, Q=5, improving by 19.9% and 6.8%, respectively, over the corresponding supervised GNN baseline. Reverse VHL-to-CRBN transfer under the same protocol yielded AUROC values of 0.702 (K=2, Q=3) and 0.821 (K=2, Q=5), confirming bidirectional applicability while revealing direction and data-regime dependence. Together, these results support ProMeta as a practical framework for cross-ligase few-shot prediction under the evaluated support/query protocols.
Comments13 pages, 4 figures, and 2 tables. Source code and reproducibility resources are available at https://github.com/yeyufeiyyf/prometa and https://doi.org/10.5281/zenodo.21371599