AI 中文总结
研究针对病原体基因组监测瓶颈从数据生成转向分析的问题,提出BioSecBench-Surveillance基准测试,含100项评估,通过提供人类分析师的数据和背景来测试AI代理,给出不同模型配置的测试结果,为衡量AI代理能否胜任基因组监测提供标准。
AI 中文摘要
随着病原体基因组监测规模的扩大,瓶颈正从数据生成转向分析。我们提出了生物安全基准测试 - 监测,这是一个包含100项评估的可验证基准,用于测试人工智能代理是否能从原始测序数据和监测背景中推断出正确的分析流程。每次评估仅向代理提供人类分析师所拥有的数据和背景,然后确定性地对其结构化答案进行评分。任务涵盖七个类别,跨越不同样本类型和测序技术。在来自16个模型 - 工具对的3962次可评分尝试中,最强配置仅通过了约一半。opus 4.8与PI以50.2%领先,在83次评估中的95%置信区间为40.1%至60.3%,与GPT - 5.5与Codex并列,置信区间为40.8%至59.6%,其次是opus 4.7与PI为49.6%,置信区间为40.0%至59.2%,Sonnet 4.6与PI为48.6%,置信区间为38.9%至58.3%。即使代理调用了正确的工作流程,其错误也源于周围的选择,如应用哪些参考、阈值、过滤器和归一化。生物安全基准测试 - 监测为衡量在下一次疫情爆发时代理是否可被信任进行基因组监测提供了标准。
英文摘要
As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.