递归自改进人工智能的进化安全性:分类、风险发现与评估
Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对递归自改进AI的安全问题,提出进化安全性视角,通过分类法和风险发现评估方法,为维持安全保证提供治理原则。
中文摘要 AI 辅助
人工智能正在快速发展,能力日益增强的系统在推理、决策、科学发现和自主开发中扮演着越来越重要的角色。随着人工智能开始参与自身的改进,从模型训练和经验积累到智能体进化和自动化AI开发,递归自改进(RSI)的前景正变得越来越相关。这一转变引发了一个基本的安全问题:当系统、其积累的经验,甚至产生其继任者的过程持续变化时,如何保持安全性?我们引入进化安全性作为一种在持续和递归自改进下研究安全性的视角。它不仅关注AI系统在特定时刻是否安全,还关注安全属性如何在进化过程中变化、持久化、积累和传播。我们描述了反复出现的表现形式,包括意图漂移、错误积累、经验污染、安全属性侵蚀、评估器漂移和风险传播。然后,我们开发了一个分类法,涵盖持久智能体状态、模型状态、评估和环境反馈、计算基质以及元级更新机制。基于这一分类法,我们研究了如何跨状态、更新、轨迹和谱系发现和评估进化风险,并推导出关于修改、选择、授权、溯源和恢复的治理原则。最后,我们概述了在AI系统变得越来越持久、自适应和递归自改进时,维持安全保证的开放问题。项目资源和提议的评估系统可在该https URL获取。
英文摘要
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience accumulation to agent evolution and automated AI development, the prospect of recursive self-improvement (RSI) is becoming increasingly relevant. This transition raises a fundamental safety question: how can safety be maintained when the system, its accumulated experience, and even the process producing its successors continue to change? We introduce Evolutionary Safety as a perspective for studying safety under persistent and recursive self-improvement. It concerns not only whether an AI system is safe at a particular moment, but how safety properties change, persist, accumulate, and propagate throughout evolution. We characterize recurring manifestations, including intent drift, error accumulation, experience contamination, safety-property erosion, evaluator drift, and risk propagation. We then develop a taxonomy spanning persistent agent state, model state, evaluation and environmental feedback, computational substrate, and meta-level update mechanisms. Building on this taxonomy, we examine how evolutionary risks can be discovered and evaluated across states, updates, trajectories, and lineages, and derive governance principles for modification, selection, authorization, provenance, and recovery. Finally, we outline open problems toward maintaining safety guarantees as AI systems become increasingly persistent, adaptive, and recursively self-improving. Project resources and proposed evaluation systems are available at https://chaunceykung.github.io/evolutionary-safety-rsi.