发表机构
UET Lahore; Jiangsu University(拉合尔工程技术大学; 江苏大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过诊断多标签EC预测中的默认阈值,揭示准确性悖论:高准确率掩盖低宏F1和召回率,提出目标特异性阈值优化和共形校准作为必要后处理保障。
AI 中文摘要
酶委员会(EC)编号的自动预测在功能注释和计算药物发现中发挥着核心作用。然而,标准的多标签机器学习流程通常依赖默认决策阈值(t=0.50),假设各目标头之间的先验分布均衡。在本研究中,我们对未校准的固定决策边界在严重类别不平衡条件下进行了系统性经验诊断,涉及N=14,096个注释化合物,分为六个主要EC类别(EC1-EC6)。我们的结果凸显了一个显著的准确性悖论:虽然多标签系统达到了看似较高的77.16%平均准确率,但宏F1分数(0.3976)和宏召回率(0.3872)揭示了严重的预测崩溃。多数目标类别存在过度敏感和过度预测问题,而少数类别则表现出急剧的召回率衰减,最终导致EC6的决策边界完全崩溃(召回率=0.00%),尽管其底层判别能力尚存(ROC-AUC=0.5857)。特征相关性分析进一步揭示了拓扑指数相对于指纹密度指标存在较高的线性冗余。最终,本诊断研究表明,标准点预测掩盖了生物信息学工作流中的关键错误。我们确定目标特异性阈值优化和事后共形校准是可靠的应用机器学习和深度学习架构所必需的开源后处理保障措施。
英文摘要
Automated prediction of Enzyme Commission (EC) numbers plays a central role in functional annotation and computational drug discovery. However, standard multi-label machine learning pipelines frequently rely on default decision thresholds (t=0.50), assuming balanced prior distributions across target heads. In this study, we present a systematic empirical diagnostic of uncalibrated fixed decision boundaries operating under severe class imbalance across N = 14,096 annotated compounds categorized into six primary EC classes (EC1-EC6). Our results highlight a pronounced Accuracy Paradox: while the multi-label system achieves a deceivingly high mean accuracy of 77.16%, the macro F1-score (0.3976) and macro recall (0.3872) reveal severe predictive breakdown. Majority target classes suffer from hyper-sensitivity and over-prediction, whereas minority classes exhibit sharp recall decay, culminating in a total decision boundary collapse for EC6 (Recall = 0.00%) despite underlying discriminative power (ROC-AUC = 0.5857). Feature correlation analysis further reveals high linear redundancy among topological indices relative to fingerprint density metrics. Ultimately, this diagnostic study demonstrates that standard point predictions mask critical errors in bioinformatics workflows. We establish target-specific threshold optimization and post-hoc conformal calibration as essential, open-source post-processing safeguards for reliable applied machine learning and deep learning architectures.
Comments11 pages, 4 figures,Enzyme Commission prediction, multi-label classification, class imbalance, threshold optimization, conformal calibration, accuracy paradox, bio-data, applied machine learning, bioinformatics. imbalanced bio-data, model diagnostic, performance metrics, open-source pipeline