arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12200cs.AIcs.CRcs.CY

前沿语言模型中用于CBRN提升评估的阈值超越框架

A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models

Rahul Gupta, Abhinav Mohanty, Payal Motwani, Venkatesh Saligrama, Satyapriya Krishna, Connor Harris, Gary Anthony Ackerman, Brandon Behlendorf, Tom Hobson, Theo… 展开作者

Rahul Gupta, Abhinav Mohanty, Payal Motwani, Venkatesh Saligrama, Satyapriya Krishna, Connor Harris, Gary Anthony Ackerman, Brandon Behlendorf, Tom Hobson, Theodore Wilson, Spyros Matsoukas

首次发表
浏览论文内容

中文总结 AI 辅助

研究前沿语言模型中CBRN提升评估问题,引入阈值超越标准(TEC)框架,将提升研究分解为独立组件,通过大规模实证研究确定两种提升形式,揭示领域异质性,为相关决策提供信息并总结方法经验。

中文摘要 AI 辅助

随着前沿语言模型的发展,政策制定者和模型开发者需要评估模型接入是否相对于仅使用公共工具实质性提高非专家行为者策划高后果化学、生物、放射或核(CBRN)滥用的能力的方法。现有CBRN评估在非专家定义、威胁范围、基线、评分标准和决策规则等方面存在差异,导致研究结果难以比较。我们引入了阈值超越标准(TEC)框架,将提升研究分解为可独立执行的组件:确定非专家参与者资格、定义研究的CBRN威胁范围以及统计估计实质性提升。然后,我们在一项大规模实证研究中实施TEC框架,该研究确定了两种提升形式:生成式(模型从零开始协助计划创建)和修正主义式(模型协助现有计划的完善)。研究生成了跨CBRN领域的攻击计划,并通过专家评审进行评估以估计生成式和修正主义式提升。应用该框架,我们的实证研究揭示了领域异质性:在这种受控的预发布评估中,模型辅助计划有时获得与专家相当的指导评级,但确认的实质性提升仅限于放射领域。这些发现为缓解和部署治理决策提供了信息,而非描述已部署模型的行为。我们最后总结了未来CBRN提升评估的方法经验,强调预先指定的标准、明确的基线、生成式和修正主义估计的分离,以及初步筛选信号和确认风险确定之间的仔细区分。

英文摘要

As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Criteria (TEC) framework that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift. We then operationalize the TEC framework in a large-scale empirical study using a design that determines two forms of uplift: generative (where a model assists plan creation from scratch) and revisionist (where a model assists refinement of an existing plan). The study produced attack plans across the CBRN domains, which we evaluated through subject-matter-expert review to estimate generative and revisionist uplift. Applying the framework, our empirical study revealed domain heterogeneity: under this controlled pre-release evaluation, model-assisted plans sometimes received expert-equivalent instructional ratings, but confirmed material uplift was limited to the radiological domain. These findings informed mitigation and deployment-governance decisions rather than characterizing deployed model behavior. We conclude with methodological lessons for future CBRN uplift evaluations, emphasizing prespecified criteria, explicit baselines, separation of generative and revisionist estimates, and careful distinction between preliminary screening signals and confirmed risk determinations.

发表机构

  • Amazon(亚马逊)
  • Nemesys Insights(Nemesys洞察)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑