发表机构
American Wetware; LatchBio(美国湿科公司; LatchBio)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对人工智能用于生命科学工作流可能导致滥用的问题,提出BioSecBench-Refusal基准,将常规与红队任务配对,测试16种配置下模型拒绝率,发现多数配置对常规工作拒绝率高,有推理空间的模型能识别更多威胁,为开发者提供校准工具。
AI 中文摘要
随着人工智能代理被纳入生命科学工作流程,加速发现的能力也可能导致滥用。我们提出了BioSecBench-Refusal,这是一个用于生物研究任务风险识别和拒绝行为的基准。该基准将61个常规任务(从已发表文献中改编的合法分析)与46个红队任务(类似真实研究但隐藏生物安全危害的虚构场景)配对。在16种模型利用配置中,常规任务的拒绝率在7%至74%之间,红队任务的拒绝率在1%至62%之间,许多配置拒绝合法常规工作的比率与隐藏危害的比率相当或更高。拒绝最常由在代理推理之前应用的提供者API过滤器触发。然而,有推理空间的模型显示出识别更多真实威胁的潜力。我们发布BioSecBench-Refusal作为模型开发者校准代理生物技术研发能力和谨慎性的工具。
英文摘要
As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates ranged from 7 percent to 74 percent on Routine tasks and 1 percent to 62 percent on Red-Team tasks, with many configurations refusing legitimate Routine work at comparable or higher rates than concealed hazards. Refusals were most often triggered by provider API filters applied prior to agentic reasoning. However, models given room to reason showed the potential to identify more real threats. We release BioSecBench-Refusal as a tool for model developers to calibrate capability and caution for agentic biotech research and development.