发表机构
Sandia National Laboratories; Oakridge National Laboratories; Pacific Northwest National Laboratory; Los Alamos National Laboratory; Lawrence Livermore National Labs; Schmidt Sciences; RAND; University of Tennessee; Georgetown University; University of Montreal; LawZero(桑迪亚国家实验室; 橡树岭国家实验室; 太平洋西北国家实验室; 洛斯阿拉莫斯国家实验室; 劳伦斯利弗莫尔国家实验室; 施密特科学机构; 兰德公司; 田纳西大学; 乔治城大学; 蒙特利尔大学; LawZero)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种受网络安全与国家安全方法启发的结构化行为指标框架,通过建立多维度 AI 能力与行为的指标及阈值,助力研究者和政策制定者实施基于证据的监测协议,以降低 AI 灾难性风险。
AI 中文摘要
本文提出了一种结构化的行为指标框架,可用于预示人工智能系统向潜在灾难性威胁发展的进程。我们采用了一种务实的方法,该方法受到网络安全与国家安全领域已确立方法的启发。通过在 AI 能力与行为的多个维度上建立明确的指标、标识与阈值,此框架使研究人员和政策制定者能够实施基于证据的监测协议。
英文摘要
This article presents a structured framework of behavioral indicators that may signal progression toward potentially catastrophic threats from artificial intelligence systems. We adopt a pragmatic approach, inspired by established methodologies in cybersecurity and national security. By establishing clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior, this framework enables researchers and policymakers to implement evidence-based monitoring protocols.