arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对齐机制中的盲区:量化大语言模型的生物安全风险

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

Shu Quan, Tianfang Hao, Sitong Fang, He Geng, Jiayi Zhou, Boyuan Chen, Kaile Wang, Donghai Hong, Juntao Dai, Yaodong Yang, Jiaming Ji

arXiv 2608.02684首次发表:更新:

发表机构

Institute for Artificial Intelligence, Peking University; The Hong Kong University of Science and Technology (Guangzhou)(北京大学人工智能研究院; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对大语言模型生物安全评估盲区,构建SPIKE-Bench评估框架,发现多数模型易生成毒素序列,提出BioSafe-Guard分类器降低功能风险并开源相关资源。

AI 中文摘要

大语言模型(LLMs)正在推动生物学研究加速发展,但这一能力同时带来了关键的生物安全威胁:协助蛋白质工程的模型同样可能被提示生成预测的毒素样序列,从而降低生物误用的门槛。然而,当前的安全评估仅针对自然语言,无法判断模型生成的氨基酸序列是生物学胡话还是计算风险信号。为解决这一评估盲区,我们推出SPIKE-Bench,将7个功能类别下的631个精心整理的毒素设计提示,与SPIKE漏斗(一个通过合规性、生物学合理性和预测毒性三个阶段过滤输出的协议)相结合,生成阶段级诊断结果和聚合的功能感知指标:功能危害率(FHR)。对32个LLMs的审计显示,大多数模型会自由遵从毒素设计请求;FHR主要由生物生成能力而非安全对齐驱动,达到50.7%;弃权(不执行)率无法预测功能风险。作为缓解的第一步,我们提供BioSafe-Guard,一个领域专用分类器,在保留良性效用的同时大幅降低预测功能风险。我们在该httpsURL发布SPIKE-Bench和BioSafe-Guard,以支持对LLMs更严格的生物安全评估。

英文摘要

Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially lowering the barrier to biological misuse. Current safety evaluations, however, operate in natural language and cannot determine whether a model-generated amino acid sequence is biological gibberish or a computational risk signal. To address this evaluation blind spot, we introduce SPIKE-Bench, coupling 631 curated toxin-design prompts across seven functional categories with the SPIKE funnel, a three-stage protocol that filters output through compliance, biological plausibility, and predicted toxicity, producing stage-level diagnostics and an aggregate function-aware metric: the Functional Harmfulness Rate (FHR). An audit of 32 LLMs reveals that most models freely comply with toxin-design requests; FHR is driven primarily by biological generation capability rather than safety alignment, reaching 50.7%; and Refusal Rate fails to predict functional risk. As a first step toward mitigation, we provide BioSafe-Guard, a domain-specialized classifier that substantially reduces predicted functional risk while preserving benign utility. We release SPIKE-Bench and BioSafe-Guard at https://github.com/PKU-Alignment/SPIKE-Bench to support more rigorous biosecurity evaluation of LLMs.

CommentsAccepted to COLM 2026. 40 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑