发表机构
SecureBio(SecureBio)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 ABLE 基准,评估 LLM 智能体在蛋白质设计中使用生物 AI 工具的能力,发现当前模型虽能降低设计门槛,但在规划与知识整合上仍不稳定。
AI 中文摘要
我们引入了 ABLE,一个用于评估 LLM 智能体在双重用途蛋白质设计工作流中使用生物人工智能模型(BAIMs,如 ProteinMPNN 和 AlphaFold3)能力的基准。ABLE 通过一组涵盖结构检索、序列生成和设计验证的任务来评估智能体性能。我们评估了 15 个前沿模型,发现其中七个拒绝执行所有任务,而其余模型表现出显著的性能差异。Claude Sonnet 4 和 Gemini 3 Pro 在信息检索、工具选择和工具使用方面取得了最高分。我们进一步将模型在部分任务上的性能与专家人类基线进行了比较。我们的结果表明,当前的 LLM 能够显著降低蛋白质设计的门槛,但在规划、策略生成以及将生物学知识与工具使用相结合方面仍存在不一致性。
英文摘要
We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but remain inconsistent in planning, strategy generation, and integrating biological knowledge with tool use.
CommentsLast revised in February 2026; presented without further revision. An earlier revision at was presented at the NeurIPS 2025 Workshop on Biosecurity Safeguards for Generative AI; see https://openreview.net/forum?id=fDysOrWaGd