发表机构
Polish-Japanese Academy of Information Technology; NIS(华沙波兰-日本信息技术学院; NIS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出失效模式框架,分析高级AI导致人类灭绝或文明崩溃的多种路径,强调风险无需意识或敌意,并指出关键影响因素与防御性研究方向。
AI 中文摘要
本文开发了一个失效模式框架,用于分析高级人工智能如何可能导致人类灭绝、不可逆的文明崩溃或人类永久性失能。核心论点是,灾难性人工智能风险并不需要意识、敌意或明确的伤害人类意图。相反,风险可能通过几种不同但相互作用的路径产生,包括自主错位、有害的人类使用、组织失败和竞争性部署。这些路径的严重性取决于能力、自主性、外部访问、持久性、制度保障和恢复能力的保持等因素。分析刻意是非操作性的:它识别因果条件、经验上可处理的中间量以及防御性研究问题,而不是造成伤害的程序。
英文摘要
This article develops a failure-mode framework for analyzing how advanced artificial intelligence could contribute to human extinction, irreversible civilizational collapse, or permanent human disempowerment. The central thesis is that catastrophic AI risk does not require consciousness, hostility, or an explicit intention to harm humanity. Instead, risk may arise through several distinct but interacting pathways, including autonomous misalignment, harmful human use, organizational failure, and competitive deployment. The severity of these pathways depends on factors such as capability, autonomy, external access, persistence, institutional safeguards, and the preservation of recovery capacity. The analysis is deliberately non-operational: it identifies causal conditions, empirically tractable intermediate quantities, and defensive research questions rather than procedures for causing harm.
Comments83 pages, 20 sections, 35 references