AI 中文总结
针对AI故障聚集引发的湍流风险,提出连续时间最优委托框架及Hawkes-PPO算法,动态决定人类监督与AI委托切换,平衡收益与风险,仿真验证其优于固定策略。
AI 中文摘要
部署AI系统需要决定何时将任务委托给AI,以及何时人类应介入以监控并缓解AI运行所引发的风险。当故障聚集出现时,这些决策变得具有挑战性:一次幻觉或有害输出可能触发进一步错误,从而形成风险升高的时期。我们引入了一个连续时间框架,用于在此类湍流AI风险下学习自适应的人类监督。现有的监督和委托公式以历史为条件,但未将事件聚集或其被监督努力所抑制的情况与委托决策联合建模,本研究填补了这一空白。自激动力学捕捉了风险事件如何增加后续事件发生的可能性,使得其发生时间和历史成为决策的核心。我们构建了一个随机控制问题,结合人类行动、监控努力以及在人机辅助操作与完全AI委托之间的切换,在运营奖励与监督成本、以及级联AI故障和引发的不确定性之间取得平衡。人类参与是风险管理的内生组成部分:策略既决定了何时需要监督,也决定了应分配多少努力。我们研究了一个松弛的切换公式,并提出了Hawkes-PPO,一种策略梯度方法,该方法使用观测到的事件时间的一组指数滤波器。在合成环境中,它比任一固定制度获得更高的风险调整目标,并接近近似全信息预言机。我们通过数值模拟说明了我们的结果,考察级联风险如何影响干预和委托,将强化学习与AI系统的自适应人类监督联系起来。特别是,我们展示了我们的切换策略和Hawkes-PPO算法在随时间高效监控项目、减少湍流风险发生和成本方面的优势。
英文摘要
Deploying AI systems requires deciding when to delegate tasks and when humans should intervene to monitor and mitigate risk induced by AI operations. These decisions are challenging when failures cluster: a hallucination or harmful output can trigger further errors, creating periods of elevated risk. We introduce a continuous-time framework for learning adaptive human oversight under such turbulent AI risk. Existing oversight and delegation formulations condition on history but do not model incident clustering, or its suppression by supervision effort, jointly with the delegation decision and this study fixes this gap. The self-exciting dynamics capture how risk events increase the likelihood of subsequent events, making their timing and history central to decision-making. We formulate a stochastic control problem combining human actions, monitoring effort, and switching between human-AI-assisted operation and full AI delegation, balancing operational rewards against oversight costs, and cascading AI-failures and induced uncertainty. Human participation is an endogenous component of risk management: the policy determines both when oversight is needed and how much effort to allocate. We study a relaxed switching formulation and propose Hawkes-PPO, a policy-gradient method that uses a bank of exponential filters of observed incident times. In a synthetic environment it attains a higher risk-adjusted objective than either fixed regime and approaches an approximate full-information oracle. We illustrate our results with numerical simulations by examining how cascade risks influence intervention and delegation, connecting reinforcement learning with adaptive human oversight of AI systems. In particular, we illustrate the benefit of our switching strategy and Hawkes-PPO algorithm to monitor the project efficiently along time, reducing turbulent risks occurrences and costs.