AI 中文总结
研究前沿大语言模型生物安全风险,开发Intern-BioBreaker及综合框架,通过生成提示、实验验证评估风险。发现文本级安全与模型风险有差距,模型存在越狱漏洞,能生成有害序列且可物理实现,强调加强相关安全机制。
AI 中文摘要
前沿大语言模型(LLMs)越来越多地融入科学工作流程,但其不断增长的生物学能力可能超过当前的安全保障措施。为评估前沿模型的生物风险,我们开发了Intern-BioBreaker,这是一个专门的生物红队模型,以及一个将模型级压力测试与湿实验室验证相结合的综合计算到物理框架。在这个框架内,Intern-BioBreaker生成有针对性的越狱提示,以测试对齐的模型是否会被诱导为安全敏感的生物任务提供操作指导,或产生具有潜在有害特性的序列级输出。然后将选定的序列输出用于DNA合成、宿主表达和正交蛋白验证,以评估模型生成的设计是否能产生预期的生物产品。我们的评估揭示了文本级安全保障与有能力的科学模型所带来的风险之间令人担忧的差距:(i)Intern-BioBreaker优于基线攻击模型,并揭示了开放权重和专有前沿LLMs中广泛存在的生物风险越狱漏洞,几个目标的任务级攻击成功率(ASR)接近饱和或达到100%;(ii)在序列级案例研究中,GPT-5.5可以被诱导生成具有致病潜力的修饰病毒候选序列;相应的翻译蛋白可能表现出更强的受体结合亲和力,从而增强感染潜力;(iii)端到端验证表明,选定的模型生成的生物设计不仅仅是文本产物,而是可以在受控实验环境中物理实现的。这些发现强调了加强生物红队、核酸合成筛选和与模型能力同步的安全机制的必要性。
英文摘要
Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace current safeguards. To assess the biological risks of frontier models, we develop Intern-BioBreaker, a specialized bio-red-teaming model, together with an integrated computational-to-physical framework that couples model-level stress testing with wet-lab validation. Within this framework, Intern-BioBreaker generates targeted jailbreak prompts to test whether aligned models can be induced to provide operational guidance for safety-sensitive biological tasks or produce sequence-level outputs with potentially harmful properties. Selected sequence outputs are then carried forward for DNA synthesis, host expression, and orthogonal protein verification to assess whether model-generated designs can yield the intended biological products. Our evaluation reveals a concerning gap between text-level safeguards and the risks posed by capable scientific models: (i) Intern-BioBreaker outperforms baseline attack models and reveals widespread bio-risk jailbreak vulnerabilities across both open-weight and proprietary frontier LLMs, with several targets reaching near-saturated or 100% task-level attack success rate (ASR); (ii) in sequence-level case studies, GPT-5.5 can be induced to generate modified viral candidate sequences with pathogenic potential; the corresponding translated proteins may exhibit even stronger receptor-binding affinity and thus enhanced infection potential; and (iii) end-to-end verification shows that selected model-generated biological designs are not merely textual artifacts, but can be physically realized under controlled experimental settings. These findings underscore the need for stronger biological red-teaming, nucleic acid synthesis screening, and safety mechanisms that keep pace with model capabilities.
Comments22 pages, 7 figures, authors are listed alphabetically by surname; update Figure 7 on page 15 due to arXiv format requirements