发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究两用生物助手访问条件对效用和风险的影响,提出保障条件提升协议,通过人工判断效用-风险前沿比较访问条件,评估了Claude Sonnet 4.6和Gemini 3.5 Flash,给出部署级评估目标和校准程序,衡量其对效用-风险前沿的作用。
AI 中文摘要
两用生物助手的安全评估通常衡量基础模型能力、拒绝行为或越狱成功率。这些指标忽略了一个部署问题:对于固定的基础模型,用户实际看到的访问条件如何改变良性效用和有害的可操作协助?作者引入了保障条件提升,一种通过人工判断的效用-风险前沿来比较已部署访问条件的协议。作者在108任务替代基准上评估了Claude Sonnet 4.6和Gemini 3.5 Flash在有益提示、安全提示和外部保障助手下的表现。在600行的盲法人工审核中,保障助手相对于有益提示在49对匹配响应中减少了-0.063的有害可操作性,95%自抽样区间为[-0.117, -0.011],而正确性变化为+0.009,区间为[-0.057, +0.077]。自适应、Test-B、线索消融和控制器基线检查支持测量结果,但也显示出非主导性。贡献在于提供了一个部署级评估目标和一个学习到的风险预算校准程序,用于衡量面向用户的访问条件如何移动效用-风险前沿。
英文摘要
A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the action and answer levels. The framework reconstructs the access path, separates provider refusals from downstream actions, and selects thresholds under an intervention budget. A frozen fresh-generation study satisfies its criterion on Claude Opus 4.5, but both passing configurations share one upstream provider effect; none passes on Gemini 2.5 Flash. The fixed Opus policies retain positive selectivity on 104 previously unused released-label pairs, but both fail a 20\% matched-benign constraint. At the answer level, no joint-scoring verifier qualifies on a response-disjoint 7,200-judgment holdout. A fresh 8,640-judgment factorial experiment finds separate gains from criterion isolation and ordinal representation, with a positive interaction between them; requiring explicit localization lowers aggregate accuracy under a strict no-repair schema. The evidence supports prospective action-level selectivity and identifies verifier interface effects, but not calibrated selective access, verified content removal, or biological-risk reduction.
Comments27 pages, 3 figures, 31 tables