发表机构
University of Toronto; McMaster University(多伦多大学; 麦克马斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统评估了4比特量化与QLoRA在多种蛋白质语言模型上的效果,证明其能大幅降低内存需求并保持性能,为内存受限场景提供高效适配策略。
AI 中文摘要
背景:蛋白质语言模型(PLMs)越来越多地用于序列生成和性质预测,但其规模使得微调和部署成本高昂。量化与参数高效微调对性能、表征和生成的影响仍未得到充分刻画。结果:我们在ESM-2、ESMC、ProtBERT、ProtT5、Ankh、Ankh3和Profluent-E1上评估了4比特量化与低秩适配器微调(QLoRA)。在蛋白质预测任务中,许多模型-任务组合保留了全量微调性能的90%以上。对于最大的模型,峰值GPU内存节省接近90%,尽管性能和效率因模型、数据集和训练配置而异。QLoRA通常保留早期层表征,同时在中层和后期层诱导任务特异性适配,类似于全量微调但表征变化更小。训练速度和功耗影响更为多样。对于使用ProLLaMA、ProtGPT2、ProGen2、ProteinGLM和ESM3进行的无条件生成,4比特量化在很大程度上保留了预测的结构和序列级性质,但token级分析揭示了自回归输出分布中依赖模型的偏移。结论:QLoRA和4比特量化降低了PLM的计算需求,尤其是GPU内存使用。我们的结果支持QLoRA作为内存受限适配的首选策略,将全量微调保留给具有挑战性的任务、不稳定的架构或低验证恢复率的情况。对于生成式PLM,序列级和结构指标应辅以分布分析,因为仅靠下游预测可能遗漏量化引起的偏移。这些方法可以拓宽大规模蛋白质建模的可及性,同时需要模型和任务特定的验证。
英文摘要
Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized. Results: We evaluated 4-bit quantization and low-rank adapter fine-tuning (QLoRA) across ESM-2, ESMC, ProtBERT, ProtT5, Ankh, Ankh3 and Profluent-E1. Across protein prediction tasks, many model-task pairs retained more than 90% of full fine-tuning performance. Peak GPU memory savings approached 90% for the largest models, although performance and efficiency varied by model, dataset and training configuration. QLoRA often preserved early-layer representations while inducing task-specific adaptations in middle and late layers, resembling full fine-tuning with smaller representational changes. Training speed and power effects were more varied. For unconditional generation with ProLLaMA, ProtGPT2, ProGen2, ProteinGLM and ESM3, 4-bit quantization largely preserved predicted structural and sequence-level properties, but token-level analysis revealed model-dependent shifts in autoregressive output distributions. Conclusion: QLoRA and 4-bit quantization reduce PLM computational requirements, particularly GPU memory usage. Our results support QLoRA as a first-pass strategy for memory limited adaptation, reserving full fine-tuning for challenging tasks, unstable architectures or low validation recovery. For generative PLMs, sequence-level and structural metrics should be complemented with distributional analysis, since downstream predictions alone may miss quantization-induced shifts. These approaches can broaden access to large-scale protein modelling while requiring model- and task-specific validation.