AI 中文总结
本研究提出含六大属性的威胁行为者画像分类学,作为开放权重前沿模型预发布风险管理的研究基础设施,以明确对手特征,支撑可解释、可比且忠实于目标风险的评估。
AI 中文摘要
前沿AI滥用风险的预发布风险管理中,通常将威胁行为者的假设隐含处理,或存在不一致、缺乏依据的问题。本研究认为,明确的对手特征刻画应被视为评估的前提,这类评估需具备可解释性、可比性,且忠实于其针对的风险。我们提出一个包含六个属性的分类学,涵盖技术复杂度、先验领域知识、组织能力、运营基础设施、财务能力和时间跨度,其基于现有恐怖主义、生物安全和网络安全文献得出的经验性层级构建。该分类学旨在作为研究基础设施,成为评估开展前预先指定对手假设的通用语言,类似医学和经济学中随机对照试验(RCT)的预分析计划。对于开放权重模型开发者而言,其应用尤为紧迫,因为这类模型的发布决策不可逆转,必须预判对手的推理。
英文摘要
Pre-release risk management for frontier AI misuse risks routinely leaves threat actor assumptions implicit, inconsistently specified, or ungrounded. This capstone argues that explicit adversary characterization should be regarded as a prerequisite for evaluations that are interpretable, comparable, and faithful to the risks they target. We propose a six-attribute taxonomy (covering technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity, and time horizon) with empirically grounded tiers derived from existing terrorism, biosecurity, and cybersecurity literature. The taxonomy is designed to function as research infrastructure: a common language for pre-specifying adversary assumptions before evaluations are conducted, analogous to pre-analysis plans for randomized controlled trials (RCTs) in medicine and economics. Its application is particularly urgent for open-weight model developers, for whom release decisions are irreversible and must anticipate adversarial reasoning.