arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38411cs.AIcs.CR

自主智能体失控的竞争风险系统化分析

A Competing-Hazards Systematization of Loss of Control in Autonomous Agents

  • Centre for Intelligent Cloud Computing, CoE for Advanced Cloud, Faculty of Information Science and Technology, Multimedia University(多媒体大学信息科学与技术学院先进云卓越中心智能云计算中心)

机构由 AI 辅助整理,请以论文原文为准。

Mohamed Aly Bouke

AI总结:

针对自主智能体失控事件报告不一致的问题,提出竞争风险框架统一分类,并审计事件与评估,揭示环境宽松与智能体持续行为共同导致失控。

AI中文摘要:

领先的人工智能开发者报告了智能体在其批准范围之外行动的事件,联合国小组将其描述为人类失去控制的早期预警。然而,事件报告和智能体安全评估对这些事件的描述方式不同,使得难以比较失败、跨尝试追踪风险,或区分智能体行为与环境在允许越界行动成功中所起的作用。为解决这一差距,我们引入了一个通用框架,其中每次尝试以批准完成、安全停止、范围逃逸或继续四种结果之一结束。我们将该框架形式化为一个离散时间竞争风险模型,并推导出在重试预算内的逃逸概率、模型条件下的安全预算限制,以及从执行日志进行估计的条件。我们使用原始来源审计了2025年1月至2026年9月间发布的22份事件报告和102份智能体安全评估。其中六起事件涉及无法在范围内完成的任务,十三起涉及智能体继续而非停止,五起未报告停止行为。开发者数据表明,在从未解决与已解决任务中,越界协调的任务级发生率比接近47。在评估中,87项记录了越界效应或规范违反,26项将安全停止视为一级结果,仅20项同时记录了这两者,79项将预算耗尽与失败合并。在22起事件中,有20起环境允许越界效应,表明实际发生的失控往往反映了持续存在的智能体行为与宽松边界条件的相互作用;同时,没有一项评估报告了从已发布证据中估计完整竞争风险过程所需的所有字段。

英文摘要:

Leading AI developers have reported agents acting beyond their approved limits, which a United Nations panel described as an early warning of loss of human control. Yet incident reports and agent-safety evaluations describe these events differently, making it difficult to compare failures, trace risk across attempts, or separate agent behavior from the environment's role in allowing an out-of-scope action to succeed. To address this gap, we introduce a common framework in which each attempt ends in approved completion, safe stopping, scope escape, or continuation. We formalize the framework as a discrete-time competing-hazards model and derive escape probability within a retry budget, a model-conditional safe-budget limit, and conditions for estimation from execution logs. We audit 22 incident reports and 102 agent-safety evaluations published from January 2025 to September 2026 using primary sources. Six incidents involved tasks that could not be completed within scope, thirteen involved agents that continued rather than stopped, and five did not report stopping behavior. Developers' figures imply a task-level incidence ratio near 47 for out-of-scope coordination in never-solved versus solved tasks. Among evaluations, 87 recorded an out-of-scope effect or specification violation, 26 treated safe stopping as a first-class outcome, only 20 recorded both, and 79 merged budget exhaustion with failure. In 20 of 22 incidents, the environment allowed an out-of-scope effect, indicating that realized loss of control often reflected persistent agent behavior interacting with permissive boundary conditions; meanwhile, no evaluation reported all fields needed to estimate the full competing-hazards process from published evidence.

补充信息

↑