arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24242cs.CY

智能体AI的预期性人类监督:一种哲学解释

Anticipatory Human Oversight of Agentic AI: A Philosophical Account

Kevin Baum, Maximilian Kiener, Markus Langer, Johann Laux

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出智能体AI需要预期性人类监督,通过规范行动议程补充反应性干预,并构建了相应的责任架构,以应对自主性与风险控制的张力。

中文摘要 AI 辅助

人类监督被广泛认为可以减轻AI系统的风险。即使对于在可识别的决策点产生离散输出的系统,将人类监督作为反应性措施的实现也在经验上是脆弱的,但日益被充分理解。然而,对于智能体AI——即在扩展的时间范围内规划、分解目标并执行多步骤行动的系统——反应性监督达到了其结构性极限:对单个行动的干预破坏了部署所激励的自主性,而对聚合模式的干预对于其累积后果仅在事后才变得可读的危害来说过于粗糙。本文认为,反应性监督必须辅以预期性模式:在智能体行动之前行使监督,通过指定构建允许行动空间的规范性议程,并通过规范、运行时和检查进行迭代细化。两者是互补的——议程的升级条件指定了何时调用反应性干预。借鉴有意义的人类控制,我们将预期性监督解读为远端原因追踪的操作化。此外,我们认为,所提出的框架通过设计产生了一种特定的责任架构:占据预期性模式是履行基于角色的前瞻性义务,而回溯性责任采取严格道德可回答性的形式——理性主义的、关系性的,并且无论过错如何都成立,基于委托人事先的预防机会。我们发展了桥接故障模式,解决了包括道德运气和控制幻觉在内的反对意见,并以监管、架构和经验意义作为结尾。

英文摘要

Human oversight is widely held to mitigate the risks of AI systems. Even for systems that produce discrete outputs at identifiable decision points, the realisation of human oversight as a reactive measure is empirically fragile, yet increasingly well understood. However, for agentic AI -- systems that plan, decompose goals, and execute multi-step actions over extended horizons -- reactive oversight reaches its structural limits: intervention on individual actions defeats the autonomy that motivates the deployment, while intervention on aggregate patterns is too coarse for harms whose cumulative consequences only become legible after the fact. This paper argues that reactive oversight must be complemented by an anticipatory mode: oversight exercised before the agent acts, by specifying the normative agenda that structures the space of permissible action and refining it iteratively through specification, runtime, and inspection. The two are complements -- the agenda's escalation conditions specify when reactive intervention is invoked. Drawing on Meaningful Human Control, we read anticipatory oversight as the operationalisation of distal-reason tracking. In addition, we argue that the proposed framework yields a specific responsibility architecture by design: occupying the anticipatory mode is the discharge of a role-grounded prospective obligation, and backward-looking responsibility takes the form of strict moral answerability -- rationalistic, relational, and holding regardless of fault, in virtue of the principal's prior opportunity for precaution. We develop bridging failure modes, address objections including moral luck and the illusion of control, and close with regulatory, architectural, and empirical implications

发表机构

  • German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
  • Institute for Ethics in Technology, Hamburg University of Technology (TUHH)(汉堡工业大学技术伦理研究所)
  • Oxford Internet Institute, University of Oxford(牛津大学牛津互联网研究所)
  • Department of Psychology, University of Freiburg(弗赖堡大学心理学系)

机构由 AI 辅助整理,请以论文原文为准。

↑