arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化与高效适配蛋白质语言模型的分析

Analysis of Quantized and Efficiently Adapted Protein Language Models

Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, Serena Singh, Sepehr Bayat, Keith Pardee

arXiv 2610.00665首次发表:更新:

发表机构

University of Toronto; McMaster University(多伦多大学; 麦克马斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统评估了4比特量化与QLoRA在多种蛋白质语言模型上的效果,证明其能大幅降低内存需求并保持性能,为内存受限场景提供高效适配策略。

AI 中文摘要

背景:蛋白质语言模型(PLMs)越来越多地用于序列生成和性质预测,但其规模使得微调和部署成本高昂。量化与参数高效微调对性能、表征和生成的影响仍未得到充分刻画。结果:我们在ESM-2、ESMC、ProtBERT、ProtT5、Ankh、Ankh3和Profluent-E1上评估了4比特量化与低秩适配器微调(QLoRA)。在蛋白质预测任务中,许多模型-任务组合保留了全量微调性能的90%以上。对于最大的模型,峰值GPU内存节省接近90%,尽管性能和效率因模型、数据集和训练配置而异。QLoRA通常保留早期层表征,同时在中层和后期层诱导任务特异性适配,类似于全量微调但表征变化更小。训练速度和功耗影响更为多样。对于使用ProLLaMA、ProtGPT2、ProGen2、ProteinGLM和ESM3进行的无条件生成,4比特量化在很大程度上保留了预测的结构和序列级性质,但token级分析揭示了自回归输出分布中依赖模型的偏移。结论:QLoRA和4比特量化降低了PLM的计算需求,尤其是GPU内存使用。我们的结果支持QLoRA作为内存受限适配的首选策略,将全量微调保留给具有挑战性的任务、不稳定的架构或低验证恢复率的情况。对于生成式PLM,序列级和结构指标应辅以分布分析,因为仅靠下游预测可能遗漏量化引起的偏移。这些方法可以拓宽大规模蛋白质建模的可及性,同时需要模型和任务特定的验证。

英文摘要

Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized. Results: We evaluated 4-bit quantization and low-rank adapter fine-tuning (QLoRA) across ESM-2, ESMC, ProtBERT, ProtT5, Ankh, Ankh3 and Profluent-E1. Across protein prediction tasks, many model-task pairs retained more than 90% of full fine-tuning performance. Peak GPU memory savings approached 90% for the largest models, although performance and efficiency varied by model, dataset and training configuration. QLoRA often preserved early-layer representations while inducing task-specific adaptations in middle and late layers, resembling full fine-tuning with smaller representational changes. Training speed and power effects were more varied. For unconditional generation with ProLLaMA, ProtGPT2, ProGen2, ProteinGLM and ESM3, 4-bit quantization largely preserved predicted structural and sequence-level properties, but token-level analysis revealed model-dependent shifts in autoregressive output distributions. Conclusion: QLoRA and 4-bit quantization reduce PLM computational requirements, particularly GPU memory usage. Our results support QLoRA as a first-pass strategy for memory limited adaptation, reserving full fine-tuning for challenging tasks, unstable architectures or low validation recovery. For generative PLMs, sequence-level and structural metrics should be complemented with distributional analysis, since downstream predictions alone may miss quantization-induced shifts. These approaches can broaden access to large-scale protein modelling while requiring model- and task-specific validation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑