arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TransNRank:基于Transformer的精准新抗原排序方法

TransNRank: Towards Accurate Neoantigen Ranking with Transformer

Zhiyin An, Yuenan Hou, Shumeng Duan, Yiming Zhou, Yuanting Zheng, Leming Shi

arXiv 2608.01924首次发表:更新:

AI 中文总结

本文提出基于Transformer的深度学习框架TransNRank,通过自注意力机制和感知正样本的训练目标,在NCI等数据集上将新抗原预测Top20召回率提升至53.1%,训练轮次大幅减少,为精准免疫肿瘤学提供新方法。

AI 中文摘要

个性化新抗原预测面临诸多挑战,包括正样本稀缺、实验数据噪声、严重的类别不平衡特性以及免疫原性特征的复杂性。现有方法如线性回归和XGBoost无法建模肽段特征内的长程依赖关系和上下文关联,因此新抗原阳性召回率的性能受到限制。本文提出一种基于Transformer的新型深度学习框架,命名为TransNRank。通过利用自注意力机制,该模型可同时捕捉局部和全局特征上下文,从而更准确地识别具有免疫原性的新抗原。采用一种感知正样本的训练目标来处理类别不平衡问题,为少数正样本分配更高权重。在NCI、TESLA和HiTIDE数据集上进行了大量实验,值得注意的是,TransNRank将新抗原预测的Top20召回率上限从46.9%(96个样本中召回45个)提升至53.1%(96个样本中召回51个),同时将训练轮次从200轮减少到20轮。此外,基于TransNRank分析特征贡献后发现,锚定位点突变和TCGA表达水平在新抗原预测中发挥着出乎意料的重要作用,而移除不显著特征以降低肽段输入维度并不会严重损害模型的整体性能。该范式不仅简化了预测流程,还为新抗原发现设定了新的前沿水平,对精准免疫肿瘤学具有广泛意义。

英文摘要

Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe class imbalance trait and the complex of immunogenicity features. Prior arts, such as linear regression and XGBoost fail to model long-range dependencies and contextual relationships within peptide features, therefore the performance of neoantigen positive recall rate is limited. In this paper, we present a novel deep learning framework based on Transformer, coined as TransNRank. By leveraging the self-attention mechanism, our model captures both local and global feature contexts, enabling more accurate recognition of immunogenic neoantigens. A positive-aware training objective is utilized to handle the class imbalance problem, assigning more weights to those few positive samples. Extensive experiments are performed on NCI, TESLA and HiTIDE datasets. Notably, our TransNRank can push the upper bound top 20 recall rate of neoantigen prediction from 46.9% (45 from 96) to 53.1% (51 from 96), while reducing the training epochs from 200 epochs to 20 epochs. Furthermore, we analyze the features contribution based on TransNRank and find that the mutation at anchor and TCGA expression level play an unexpected important role in neoantigen prediction, and removing insignificant features to reduce the input dimensionality of peptides does not drastically impair the overall performance of the model. Our paradigm not only streamlines the prediction pipeline but also sets a new state-of-the-art for neoantigen discovery, with broad implications for accurate immuno-oncology.

CommentsThe final authorship has not been determined and this version has not been polished

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑