arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08059q-bio.QMcs.LG

MI-PEFT:集成混合专家的参数高效微调蛋白质语言模型提升嗜酸蛋白分类

MI-PEFT: Mixture-of-Experts Integrated Parameter-Efficient Fine-Tuning Protein Language Models Improves Acidophilic Proteins Classification

Honghan Shen

首次发表
浏览论文内容

中文总结 AI 辅助

针对嗜酸蛋白鉴定耗时问题,提出MI-PEFT框架,结合LoRA与DeepSeekMoE分类头,在ESM C-600M上实现高效分类并缓解类别不平衡。

中文摘要 AI 辅助

嗜酸蛋白在强酸性条件下保持稳定和功能,对于工业生物催化、酸相关生物加工以及耐酸酶的发现具有重要意义。然而,其鉴定高度依赖于耗时的实验筛选方法。随着蛋白质序列数据库的快速增长,对准确且高效的计算鉴定方法的需求日益增强。蛋白质语言模型(PLMs)的出现显著改善了下游生物预测任务的序列表示。本文提出MI-PEFT,一种集成混合专家的参数高效微调框架。该框架基于ESM C-600M骨干,整合了基于LoRA的PEFT方法和基于DeepSeekMoE的分类头,以解决PEFT的局限性并显著提高计算效率。值得注意的是,该任务的特点是数据集中存在显著的类别不平衡,使得高特异性尤其具有挑战性。实验结果表明,MI-PEFT在PLMs上,特别是C³A,可作为识别嗜酸蛋白的高效工具,并通过保留预训练表示来解决类别不平衡的受限途径。

英文摘要

Acidophilic proteins that remain stable and functional under highly acidic conditions, are important for industrial biocatalysis, acid-related bioprocessing, and the discovery of acid-stable enzymes. However, their identification relies heavily on time-consuming experimental screening methods. With the rapid growth of protein sequence databases, the need for computational identification methods that are both accurate and efficient has become stronger. The emergence of protein language models (PLMs) has significantly improved the sequence representation of downstream biological prediction tasks. This paper proposes MI-PEFT, a mixture-of-experts integrated parameter-efficient fine-tuning framework. Built on the ESM C-600M backbone, the framework incorporates LoRA-based PEFT methods and a DeepSeekMoE-based classification head to resolve the limitations of PEFT and significantly improve computational efficiency. Notably, this task is characterized by a significant class imbalance in the dataset, making high specificity particularly challenging. The experimental results demonstrate that MI-PEFT on PLMs, especially {\text{C}}^{\text{3}}\text{A}, serves as an efficient tool for identifying acidophilic proteins and a constrained pathway that helps resolve class-imbalance by preserving the pretrained representations.

发表机构

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

↑