arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨肽和靶标偏移的肽-蛋白亲和力预测基准测试

Benchmarking Peptide-Protein Affinity Prediction Across Peptide and Target Shifts

Jiaxin Tian, Darren An, Jun Li

arXiv 2608.30175首次发表:更新:

发表机构

College of Biology, Hunan University; Lingang Laboratory(湖南大学生命科学学院; 临港实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究整合多源肽-蛋白结合数据构建基准,测试多种肽表示、ESM-2嵌入及回归器在不同划分下的性能,发现表示与回归器泛化特性,提出基准需匹配预期用途并联合评估多因素影响。

AI 中文摘要

肽-蛋白亲和力模型通常采用单一数据划分方式进行评估,这会掩盖它们是在观测靶标的测量值之间进行插值,还是能跨肽或靶标偏移进行泛化。我们整合了三个来源的定量肽-蛋白结合数据,得到11349个去重配对,并在肽相似性、靶标内和留一靶标划分下,对10种肽表示、ESM-2蛋白嵌入和6种回归器进行基准测试。在60种匹配的表示-回归器配置中,平均测试斯皮尔曼相关系数分别为0.462、0.669和0.530。前配置在肽相似性和靶标内划分下采用带随机森林的ECFP-16计数指纹,在排除精确靶标序列时切换为带Extra Trees的HELM-BERT。跨划分的表示秩相关系数范围为-0.042至0.624,而回归器秩相关系数范围为0.771至0.943。学习曲线显示,在监督有限时表示差异最大,随训练数据增加而缩小。在测试协议下,PeptideCLM-2适配和简单逐元素交互特征未比冻结编码器与直接拼接提供一致增益。这些结论适用于汇集了转换后的Kd、Ki和IC50测量值的数据集,以及精确序列级别的靶标排除。因此,肽-蛋白亲和力基准测试应使数据划分与预期用途一致,并联合评估数据规模、分子表示和下游学习器的影响。

英文摘要

Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed targets or generalize across peptide or target shifts. We integrated three sources of quantitative peptide-protein binding data to obtain 11,349 deduplicated pairs and benchmarked ten peptide representations, ESM-2 protein embeddings, and six regressors under peptide-similarity, within-target, and leave-target-out partitions. Across 60 matched representation-regressor configurations, mean test Spearman correlations were 0.462, 0.669, and 0.530, respectively. The top configuration shifted from ECFP-16 count fingerprints with random forest in the first two settings to HELM-BERT with Extra Trees when exact target sequences were excluded. Representation-rank correlations ranged from -0.042 to 0.624 across partitions, whereas regressor-rank correlations ranged from 0.771 to 0.943. Learning curves showed that representation differences were largest with limited supervision and narrowed as training data increased. PeptideCLM-2 adaptation and simple element-wise interaction features provided no consistent gain over a frozen encoder and direct concatenation under the tested protocols. These conclusions are specific to a dataset that pools transformed Kd, Ki, and IC50 measurements and to target exclusion at the exact-sequence level. Peptide-protein affinity benchmarks should therefore align data partitions with the intended use and jointly assess the effects of data scale, molecular representation, and downstream learner.

Comments11 pages, 6 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑