arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05436q-bio.BMcs.LG

新型混合蛋白质支架缺口填充:加权机器学习集成、束搜索与质量约束重排序

Novel hybrid protein scaffold gap filling using weighted machine learning ensemble, beam search, and mass-constrained reranking

Tahmid Enam Shrestha, Md. Manzurul Hasan, Md. Rafiqul Islam

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种混合机器学习与质量约束重排序框架,结合加权集成、束搜索及同源检索,在已知缺口大小和质量条件下实现高准确率的蛋白质支架缺口填充。

中文摘要 AI 辅助

蛋白质支架缺口填充是蛋白质序列重建中的一项重要计算任务,其中缺失的氨基酸区域必须从不完整的支架信息中推断出来。本研究提出了一种混合机器学习与质量约束重排序框架,用于在已知缺口大小和已知缺口质量设置下进行蛋白质支架缺口填充。使用来自MabCampath、P5A蛋白形式(proteoform)和碳酸酐酶2的同源蛋白质序列,生成了掩蔽的11-mer残基级样本和全缺口评估案例。残基预测任务被表述为一个20类氨基酸分类问题,采用首位、中间位和末位掩蔽策略。使用原始编码、行平均和SVD降维特征训练了多个经典机器学习模型,并通过验证准确率加权集成将最强模型组合起来。对于已知大小的缺口重建,使用束搜索从残基级概率估计中生成完整的缺失肽序列。对于已知质量的重建,将质量约束的同源候选检索与基于质量有效性、同源频率、上下文支持、集成似然、质量误差和长度惩罚的混合重排序相结合。所提出的框架在残基级验证准确率上达到95.41%,在已知大小缺口的精确匹配准确率上达到87.50%,并在七个CAH2已知质量基准案例中实现了100%的前5名恢复率。这些结果表明,所提出的框架通过整合局部序列学习、同源证据、肽质量约束和生化验证,能够有效重建缺失的蛋白质区域。

英文摘要

Protein scaffold gap filling is an important computational task in protein sequence reconstruction, where missing amino acid regions must be inferred from incomplete scaffold information. This study proposes a hybrid machine learning and mass constrained reranking framework for protein scaffold gap filling under known-gap-size and known-gapmass settings. Homologous protein sequences from MabCampath, P5A proteoform, and carbonic anhydrase 2 were used to generate masked 11-mer residue-level samples and fullgap evaluation cases. The residue prediction task was formulated as a 20-class amino acid classification problem using first-, middle-, and last-position masking. Multiple classical machine learning models were trained using raw encoded, row-average, and SVD-reduced features, and the strongest models were combined through a validation-accuracy-weighted ensemble. For known-size gap reconstruction, beam search was used to generate complete missing peptide sequences from residue-level probability estimates. For known-mass reconstruction, mass-constrained homologous candidate retrieval was combined with hybrid reranking based on mass validity, homologous frequency, context support, ensemble likelihood, mass error, and length penalty. The proposed framework achieved 95.41% residue-level validation accuracy, 87.50% known-size exact-match accuracy, and 100% top-5 recovery on seven CAH2 known-mass benchmark cases. These results indicate that the proposed framework can effectively reconstruct missing protein regions by integrating local sequence learning, homologous evidence, peptide mass constraints, and biochemical validation.

发表机构

  • American International University-Bangladesh(美国国际大学-孟加拉国分校)
  • City University(城市大学)

机构由 AI 辅助整理,请以论文原文为准。

↑