AI 中文总结
本研究将蛋白质注释视为序列到文本生成问题,用QLoRA微调30亿参数的Ministral 3模型,通过GPT专家评估,证明紧凑LLM可生成有生物学价值的策展式注释。
AI 中文摘要
对新测序蛋白质的功能注释仍然是分子生物学中的一个瓶颈:公共数据库中序列数量的增长远快于人工注释的能力。大多数计算方法将注释视为固定本体上的多标签分类,这限制了预测只能针对预定义的标签集。在本研究中,我们将蛋白质注释视为一个序列到文本的生成问题。我们使用QLoRA(4位NF4量化与低秩适配器)对30亿参数的Ministral 3基础模型进行微调,训练数据为序列注释对。我们采用LLM作为专家的协议来评估预测:一个被提示为资深分子生物学策展人的GPT模型,对生物体识别进行二元评分,并对功能注释质量进行评分。我们得出结论,经过QLoRA微调的紧凑型大语言模型能够为相当一部分蛋白质生成具有真实生物学价值的策展式注释。我们还讨论了数据质量、模型扩展和证据基础方面的未来方向,这些对于使该方法在实际应用中足够可靠是必要的。
英文摘要
Functional annotation of newly sequenced proteins remains a bottleneck in molecular biology: the number of sequences in public repositories grows far faster than the capacity for manual curation. Most computational approaches consider annotation as multi-label classification over a fixed ontology, which constrains predictions to a predefined label set. In this work we study the the protein annotation as a sequence-to-text generation problem. We fine-tune the 3B-parameter Ministral 3 base model with QLoRA (4-bit NF4 quantization with low-rank adapters) on sequence annotation pairs. We assess predictions with an LLM-as-expert protocol: a GPT model prompted as a senior molecular-biology curator scores organism identification as binary and function annotation quality. We conclude that QLoRA-fine-tuned compact LLMs can generate curator-style annotations with genuine biological value for a substantial subset of proteins. We also discuss future directions in data quality, model scaling, and evidence grounding that are needed to make the approach sufficiently reliable for practical use.