arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28335cs.AI

面向欧盟《人工智能法案》实践准则的系统性风险证据的开放管道与仪表盘

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出系统性风险指数,一个开放评估管道与仪表盘,将19个基准映射到欧盟实践准则的四类风险,通过最坏情况聚合揭示平均评估隐藏的风险,并验证了LLM评判与人类评分的一致性。

中文摘要 AI 辅助

关于人工智能安全的主张所影响的受众远远超出人工智能社区,然而许多主张依赖于不透明的证据或静态评估,甚至有时根本无法获取支持性证据。我们提出了系统性风险指数(Systemic Risk Index),这是一个开放的评估管道和仪表盘,旨在让经验证据对公众更加透明和可追溯。我们的工作将19个公开基准整理为欧盟通用人工智能实践准则(EU GPAI Code of Practice)所定义的四个系统性风险类别——CBRN(化学、生物、放射性和核)、网络攻击、有害操纵以及失控——并使用保留危害的扰动和模拟部署情境来评估模型。交互式仪表盘允许用户在平均聚合和最坏情况聚合之间切换,调整模型能力对聚合分数的影响方式,并将每个风险评级追溯到其基准证据。在18个模型中,在最坏情况聚合下,分数下降了14至37分,这凸显了在模型风险的平均评估中可能被隐藏的信息。法学硕士(LLM)评判员与人类评分者的一致性接近人类与人类之间的一致性(κ=0.78–0.82),而一项盲审发现,83%的抽样变换保留了原始危害。在一项调查(N=21)中,大多数参与者报告称分数易于理解,并且仪表盘鼓励他们在不同设置下查看模型评估。

英文摘要

Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($κ= 0.78\text{--}0.82$), and a blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings

发表机构

  • UC San Diego(加州大学圣迭戈分校)
  • University of Maryland(马里兰大学)
  • Vector Institute(向量研究所)
  • EuroSafeAI
  • Purdue University(普渡大学)
  • Brown University(布朗大学)
  • ETH Zurich(苏黎世联邦理工学院)
  • MPI for Intelligent Systems, Tübingen(马克斯·普朗克智能系统研究所(蒂宾根))
  • University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑