发表机构
Indian Institute of Technology Roorkee; Indian Institute of Science(印度鲁基工业学院; 印度科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种仅用序列描述符结合TabPFN的简单流程,在ESCAPE数据集上多活性抗菌肽预测性能超越主流复杂模型,且对远程同源物提升显著,推理无需结构预测。
AI 中文摘要
抗菌肽(AMPs)常针对多种病原体类别发挥作用,因此多标签活性预测比二元抗菌分类更符合实际筛选目标。ESCAPE基准将此场景形式化,但主流方法通常依赖多模态、结构条件的深度模型,训练与调优成本高昂。本文展示了一种仅基于序列的简单流程,通过将330种可解释序列描述符与TabPFN(一种表格式基础模型,可在单次前向传播中执行上下文预测,无需基于梯度的训练或超参数搜索)相结合,可达到并超越上述方法的性能。在包含82359条肽、5个标签的ESCAPE数据集上,标签幂集TabPFN模型的mAP-5达77.8%,优于此前报告的最佳值72.1%;概率分类器链是首个同时在5个标签上达到或超越已发表最佳平均精度的方法。在先前最先进的单折训练协议下,性能提升依然存在,表明其并非训练集大小的人为结果,且对远程同源物(序列一致性低于30%)的提升最大(+11.2个百分点)。消融实验进一步表明,推理时无需预测结构,且性能并非由任何单一描述符族驱动:10个全局物理化学标量可恢复完整特征性能的91%。最后,显式建模标签依赖对稀缺活性产生针对性益处,并支持从部分正证据中对下一个待测定活性进行排序。
英文摘要
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE benchmark formalizes this setting, but leading approaches typically rely on multimodal, structure-conditioned deep models that are costly to train and tune. We show that a simple, sequence-only pipeline can match and surpass these methods by combining 330 interpretable sequence descriptors with TabPFN, a tabular foundation model that performs in-context prediction in a single forward pass without gradient-based training or hyperparameter search. On ESCAPE (82,359 peptides; five labels), a label-powerset TabPFN model achieves mAP-5 = 77.8%, improving on the previously best reported 72.1%. A probabilistic classifier chain is the first method to match or exceed the best published average precision on each of the five labels simultaneously. The gains persist under the prior state-of-the-art single-fold training protocol, indicating they are not a training-set-size artefact, and are largest for remote homologues (+11.2 points below 30% sequence identity). Ablations further show that predicted structure is unnecessary at inference and that performance is not driven by any single descriptor family: ten global physicochemical scalars recover 91% of full-feature performance. Finally, explicitly modelling label dependence yields targeted benefits for scarce activities and supports ranking which activity to assay next from partial positive evidence.