arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03178cs.CY

AI风险建模中的开放问题:来自AI风险建模技术基础研讨会的见解

Open Problems in AI Risk Modeling: Insights from a Workshop on the Technical Foundations of AI Risk Modeling

Krystal Jackson, Deepika Raman, Jakub Kryś, Sean P. Fillingham, Jack Kengott, Andrew J. Lohn, Nada Madkour, Henry Papadatos, James Sykes, Anna Katariina Wisakanto, Malcolm Murray

首次发表
浏览论文内容

中文总结 AI 辅助

本文梳理AI风险建模领域的开放问题,对比主流方案,结合专家研讨明确待解方向,提出需整合定量建模与独立评估等机制推进该领域发展。

中文摘要 AI 辅助

我们研究用于评估高级AI系统带来的社会风险的稳健风险模型设计,这是AI治理领域的新兴方向。许多监管提案日益要求进行系统性风险评估,但由于缺乏严格的定量方法,目前实践中最先进的风险建模应呈现何种形态仍是待解问题。我们梳理了该问题涉及的5种研究传统:概率风险评估、灾难性AI风险分析、网络安全风险量化、贝叶斯因果推断及基于阈值的治理;对比了两种主流方案:基于场景的风险估计与基于贝叶斯网络的阈值设定。结合22位专家参与的研讨会及后续分析,我们明确了模型结构、范围、证据整合、验证、治理方面的一系列开放问题,最后提出了进展优先方向,指出这一进展依赖于将定量建模与独立评估、透明分层披露、能长期维护和更新风险模型的机构相结合。

英文摘要

We investigate the design of robust risk models to assess societal risks posed by advanced AI systems, an emerging area in AI governance. Many regulatory proposals increasingly require systemic risk assessment, but in the absence of rigorous quantitative methods, the question remains what state of the art risk modeling should look like in practice. We identify the key methodological and institutional challenges that currently limit the adoption of risk modeling. We review five research traditions that inform this problem: probabilistic risk assessment, catastrophic AI risk analysis, cybersecurity risk quantification, Bayesian causal inference, and threshold-based governance. We compare two leading proposals, scenario-based risk estimation and Bayesian network-based threshold setting. Drawing on a workshop with 22 experts and subsequent analysis, we identify a structured agenda of open questions concerning model structure, scope, evidence integration, validation, and governance. We close by outlining priorities for progress, arguing that it will depend on integrating quantitative modeling with independent evaluation, transparent and tiered disclosure, and institutions capable of maintaining and updating risk models over time.

发表机构

  • Institute for Security and Technology(安全与技术研究所)
  • Center for Long-Term Cybersecurity(长期网络安全中心)
  • SaferAI
  • Center for Security and Emerging Technology(新兴技术安全中心)
  • University of Warwick(华威大学)
  • Center for AI Risk Management & Alignment(人工智能风险管理与对齐中心)

机构由 AI 辅助整理,请以论文原文为准。

↑