BioFirewall:面向智能体AI设计阶段生物安全筛查的基因组写入原生治理层
BioFirewall: A genome-writing-native governance layer for design-stage biosecurity screening of agentic AI
浏览论文内容
中文总结 AI 辅助
针对智能体AI基因组编辑设计阶段的生物安全管控缺口,提出BioFirewall中间件,其在多维度筛查中表现优于大模型,能有效抵御攻击且假拒绝率低,已开源并附带可复现基准。
中文摘要 AI 辅助
背景:人工智能设计工具如今可规划基因组规模的编辑,而智能体系统在人类监督日益减少的情况下执行这些规划。生物安全控制仅局限于两个节点:基础模型的拒绝护栏和合成器的序列身份筛查。二者之间的设计阶段(即规划被指定的阶段)仍仅受建议约束,无任何已部署系统进行管控。结果:我们提出了BioFirewall,这是一种受规则约束的中间件,可拦截基因组写入规划,并针对基因组写入特有的五个危害维度( cargo、基因座、编辑类型、种系、规模)返回“允许”“标记待审核”或“拒绝”的结果,同时提供引用证据、已签名的设计通行证、防篡改审计日志和分层访问权限。在针对独立预言机评分的去循环安全代理基准测试中,功能感知型 cargo 分类器在1%的假阳性率下达到了0.72的真阳性率(95%置信区间:0.43至0.89),而前沿大模型和开源语言模型裁判无法可靠筛查相同序列。在提示注入攻击下,开源权重裁判在每个通道的6次试验中分别有3次和5次将阻止裁决翻转为允许,而确定性筛查保持不变。来自三个模板的288个合法规划均未被拒绝,得出假拒绝率的95%置信上限为0.0103,且会话监视器拦截了跨调用分解攻击。在保留的基因集上,基因座维度富集了体内插入致癌的驱动因素(AUROC为0.605;比值比为3.34)。结论:设计阶段的治理在实践中是可行的。BioFirewall以开源形式发布,附带预注册的、可通过开放数据复现的基准测试。
英文摘要
Background. Artificial-intelligence design tools now plan genome-scale edits, and agentic systems execute those plans with progressively less human oversight. Biosecurity controls are limited to two points: refusal guardrails at the foundation model and sequence-identity screening at the synthesiser. The design stage between them, where the plan is specified, remains governed by recommendations rather than any deployed system. Results. We present BioFirewall, a rule-governed middleware that intercepts a genome-writing plan and returns allow, flag-for-review, or refuse across five hazard axes native to genome writing: cargo, locus, edit type, germline and scale, with cited evidence, a signed design passport, a tamper-evident audit log, and tiered access. On a de-circularised benchmark of safe proxies scored against independent oracles, a function-aware cargo classifier reached a true-positive rate of 0.72 (95% CI 0.43 to 0.89) at a 1% false-positive rate, whereas frontier and open language-model judges did not screen the same sequences reliably. Under prompt injection, the open-weight judges flipped their blocking verdict to allow in 3 and 5 of 6 trials per channel, while the deterministic screen remained invariant. None of 288 legitimate plans from three templates was refused, yielding a certified 95% upper bound of 0.0103 on the false-refuse rate, and a session monitor intercepted cross-call decomposition attacks. On a held-out gene set, the locus axis was enriched for drivers of in vivo insertional oncogenesis (AUROC 0.605; odds ratio 3.34). Conclusions. Design-stage governance is achievable in practice. BioFirewall is released as open source with a pre-registered, open-data-reproducible benchmark.