BERT4DTI:基于BERT的药物-蛋白质相互作用预测模型
BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions
浏览论文内容
中文总结 AI 辅助
提出BERT4DTI模型,结合ChemBERTa和ProtBERT编码及双向互注意力,以较低参数量在多个基准上高效预测药物-蛋白质相互作用。
中文摘要 AI 辅助
理解药物如何与蛋白质靶标相互作用是药物发现、药物重定位以及在昂贵的实验测试之前早期识别有前景的治疗候选药物的基础。基于序列的DTI模型面临三个实际限制:标记的相互作用数据稀缺且分布不均,大型预训练的化学和蛋白质编码器端到端微调成本高昂,以及独立编码的序列无法捕捉配对特异性依赖。我们提出BERT4DTI,它使用ChemBERTa编码SMILES字符串,使用ProtBERT编码氨基酸序列,在token级表示之间应用双向互注意力,并使用卷积层和多层感知机对生成的相互作用特征进行分类。为减少可训练参数量,ProtBERT被截断为保留18层,且每个编码器仅微调最后两层。在BIOSNAP、DAVIS和BindingDB上,BERT4DTI具有竞争力,在BIOSNAP上取得最佳ROC-AUC和PR-AUC,并在所有三个基准上取得最高灵敏度。在DAVIS上的消融实验表明,互注意力提升了PR-AUC和特异性。与完整BERT微调的353M参数相比,BERT4DTI仅需125M可训练参数,为基于序列的DTI筛选提供了有利的性能-参数权衡,而运行时性能分析、校准和泄漏审计验证留待未来工作。
英文摘要
Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce and unevenly distributed, large pretrained chemical and protein encoders are expensive to fine-tune end-to-end, and independently encoded sequences do not capture pair-specific dependencies. We present BERT4DTI, which encodes SMILES strings with ChemBERTa and amino-acid sequences with ProtBERT, applies bidirectional mutual attention between token-level representations, and classifies the resulting interaction features using convolutional layers and a multilayer perceptron. To reduce trainable size, ProtBERT is truncated to 18 retained layers and only the last two layers of each encoder are fine-tuned. On BIOSNAP, DAVIS and BindingDB, BERT4DTI is competitive, achieving the best ROC-AUC and PR-AUC on BIOSNAP and the highest sensitivity on all three benchmarks. An ablation on DAVIS shows that mutual attention improves PR-AUC and specificity. With 125M trainable parameters compared with 353M for full BERT fine-tuning, BERT4DTI provides a favourable performance-parameter trade-off for sequence-based DTI screening, while leaving runtime profiling, calibration and leakage-audited validation for future work.