arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能前沿的孤儿风险:不同的安全与合规框架揭示了AI公司如何选择其优先处理的风险

Orphan risks at the frontier of artificial intelligence: What diverging safety and compliance frameworks reveal about how AI companies choose the risks they prioritize

Andrew D. Maynard

arXiv 2608.16895首次发表:更新:

AI 中文总结

本文对比四家前沿AI公司2023-2026年的安全合规文件,识别出其风险选择的四个筛选标准,提出“安全差异”概念,揭示了孤儿风险的成因及轻量应对工具。

AI 中文摘要

开发全球最强大人工智能系统的公司在梳理其技术带来的风险方面出奇地勤勉。然而,处于新兴前沿模型与其在经济上成功且对社会有益的部署之间的风险格局正变得日益难以应对。更复杂的是,许多前沿AI公司对其技术可能出现的问题持有不止一种描述。本文通过对比Anthropic、OpenAI、Google DeepMind和Meta在2023年至2026年间发布的安全与合规文件,记录了这些描述之间的分歧,并探讨由此产生的记录揭示了这些公司如何选择其管理的风险。由于这些文件带有时间戳并已归档,它们提供了正在进行的机构风险选择的宝贵公共记录。从该记录中,本文确定了决定哪些风险倾向于在自行制定的框架中留存的四个筛选标准(可测量性、严重性、可审计性和竞争成本),并引入“安全差异”作为公司为自身选择的风险格局与监管机构为其选择的风险格局之间的差距。虽然急性、可量化的风险在各文件中均有体现,但诸如有害操纵等较难处理的风险在法律要求披露的地方被清晰阐述,却仍未出现在大多数自行选择的框架中。这种排除源于这些机构对风险的定义方式。本文借鉴机构风险选择的学术研究和风险创新框架,展示了将风险重新定义为对价值的威胁如何有助于解释风险如何成为“孤儿风险”、指出未来可能出现的意外情况,以及如何为消除孤儿风险提供轻量工具,而这些工具是前沿AI的安全机构目前未被组织起来处理的。

英文摘要

Companies developing some of the world's most powerful artificial intelligence systems are surprisingly diligent in how they map out the risks their technologies present. Yet the risk landscape that lies between emerging frontier models and their economically successful and societally beneficial deployment is becoming increasingly hard to navigate. Complicating this further, many frontier AI companies maintain more than one account of what could go wrong with their technologies. This paper documents the divergence between these accounts by comparing safety and compliance documents published by Anthropic, OpenAI, Google DeepMind and Meta between 2023 and 2026, and considers what the resulting record reveals about how these companies select the risks they manage. As these documents are timestamped and archived, they provide a valuable public record of institutional risk selection in progress. From this record the paper identifies four filters that determine which risks tend to survive in self-authored frameworks (measurability, severity, auditability and competitive cost) and introduces the "safety differential" as the gap between the risk landscape a company selects for itself, and the one regulators select for it. While acute, quantifiable risks appear across documents, less tractable risks such as harmful manipulation are articulated fluently where law compels disclosure, yet remain absent from most self-chosen frameworks. This is an exclusion that follows from how these institutions define risk. Drawing on scholarship on institutional risk selection and the framework of risk innovation, the paper shows how redefining risk as a threat to value can help explain how risks become "orphan risks," how it indicates where future blindsides may occur, and how it points to lightweight tools for de-orphaning risks that frontier AI's safety apparatuses are not currently organized to address.

Comments21 pages, 40 references. Also available on SSRN: https://dx.doi.org/10.2139/ssrn.7068898

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑