PFArena:蛋白质修饰语言模型基准测试
PFArena: Benchmarking Language Models for Protein Modification
浏览论文内容
中文总结 AI 辅助
PFArena基准通过四个任务接口评估PLMs、LLMs及智能体在蛋白质修饰中的性能,发现PLMs擅长单突变生成,LLMs和智能体擅长多突变排序,且所有模型在扩大搜索空间时面临挑战。
中文摘要 AI 辅助
蛋白质修饰需要在巨大的序列空间中进行探索,然而湿实验验证仍然存在通量低和成本高的问题。尽管包括蛋白质语言模型(PLMs)、大型语言模型(LLMs)和基于LLM的智能体在内的计算范式在蛋白质修饰中显示出潜力,但它们在现实实验决策场景中的相对有效性仍不清楚。为弥合这一差距,我们提出了PFArena,一个包含四个受控任务接口的基准,涵盖单突变体生成和多突变体排序。通过提供不同水平的突变适应度数据,PFArena反映了四种具有不同程度先前实验背景的代表性研究场景。我们使用互补指标评估了六个PLMs、六个LLMs和五个基于LLM的智能体,以衡量峰值和整体蛋白质修饰性能。我们的评估揭示,模型性能随目标特异性实验证据的可用性而系统性变化:PLMs通过利用蛋白质特异性先验在开放式单突变体生成中表现出熟练度,而LLMs和智能体在多突变体排序中表现强劲,特别是在目标特异性适应度数据可用时。然而,所有模型家族在搜索空间大小和突变深度增加时都面临根本性挑战。我们发布了代码和基准套件,以促进模型辅助蛋白质修饰的可重复研究。
英文摘要
Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision-making settings remains unclear. To bridge this gap, we introduce PFArena, a benchmark comprising four controlled task interfaces that cover single-mutant generation and multi-mutant ranking. By providing varying levels of mutation fitness data, PFArena reflects four representative research scenarios characterized by differing degrees of prior experimental context. We assess six PLMs, six LLMs, and five LLM-based agents using complementary metrics to measure both peak and overall protein modification performance. Our evaluation reveals that model performance shifts systematically with the availability of target-specific experimental evidence: PLMs demonstrate proficiency in open-ended single-mutant generation by leveraging protein-specific priors, whereas LLMs and agents perform strongly in multi-mutant ranking, particularly when target-specific fitness data are available. Nevertheless, all model families face fundamental challenges with increasing search-space size and mutation depth. We release our code and benchmark suite to facilitate reproducible research in model-assisted protein modification.