发表机构
Brown University; Centre pour la Sécurité de l’Intelligence Artificielle (CeSIA)(布朗大学; 人工智能安全中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对前沿人工智能公司能力阈值差异大的问题,开发推导协调阈值的方法,在滥用风险领域以预期危害为关键要素建模,自动化人工智能研发领域基于进展速度设阈值,扩展了相关工作并指出差距局限。
AI 中文摘要
前沿人工智能公司公布的能力阈值差异很大,这使得第三方难以核实阈值是否被突破或比较不同公司的要求。此外,没有共同的最低阈值,风险缓解可能不一致,导致安全标准可能出现逐底竞争。我们开发了一种方法来推导三个风险领域的协调阈值。对于滥用风险(网络和生物),我们将预期危害作为关键要素,并使用明确的风险建模方法,该方法考虑了风险渠道和模型发布条件。对于自动化人工智能研发,我们基于观察到的人工智能进展速度而非预期危害来提出阈值。我们的分析扩展了先前的工作,并突出了现有的实证差距和局限性。
英文摘要
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.