arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

协调人工智能安全阈值

Harmonizing AI Safety Thresholds

Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza, Markov Grey

arXiv 2607.16112首次发表:更新:

发表机构

Brown University; Centre pour la Sécurité de l’Intelligence Artificielle (CeSIA)(布朗大学; 人工智能安全中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对前沿人工智能公司能力阈值差异大的问题,开发推导协调阈值的方法,在滥用风险领域以预期危害为关键要素建模,自动化人工智能研发领域基于进展速度设阈值,扩展了相关工作并指出差距局限。

AI 中文摘要

前沿人工智能公司公布的能力阈值差异很大,这使得第三方难以核实阈值是否被突破或比较不同公司的要求。此外,没有共同的最低阈值,风险缓解可能不一致,导致安全标准可能出现逐底竞争。我们开发了一种方法来推导三个风险领域的协调阈值。对于滥用风险(网络和生物),我们将预期危害作为关键要素,并使用明确的风险建模方法,该方法考虑了风险渠道和模型发布条件。对于自动化人工智能研发,我们基于观察到的人工智能进展速度而非预期危害来提出阈值。我们的分析扩展了先前的工作,并突出了现有的实证差距和局限性。

英文摘要

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑