arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为微服务系统中风险约束干预决策的安全修复

Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems

Chengxiao Dai, Zhaokun Yan, Chenjun Lei, Qiao Li, Luyan Zhang

arXiv 2607.20005首次发表:更新:

发表机构

School of Computer Science, University of Sydney; Khoury College of Computer Sciences, Northeastern University; School of Computation, Information and Technology, Technical University of Munich; College of Business and Economics, Australian National University; School of Computer Science and Technology, Fujian Normal University(悉尼大学计算机科学学院; 美国东北大学库里计算机科学学院; 慕尼黑工业大学计算、信息与技术学院; 澳大利亚国立大学商业与经济学院; 福建师范大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究微服务系统中安全修复问题,将其转化为风险约束干预决策问题并构建约束马尔可夫决策过程,引入三维风险分解和上下文自适应人在回路门,通过实验验证该框架能降低错误修复率、提高修复成功率并减轻值班升级负载。

AI 中文摘要

在现代 IT 运维中,错误修复的成本往往超过不采取行动的成本。现有自动化修复系统旨在生成操作而非决定是否需要干预,安全仅通过人工审批事后执行。本文有三项贡献:一是将安全修复重新表述为风险约束干预决策问题并转化为约束马尔可夫决策过程,使智能体在有界错误修复率下最大化修复成功率;二是引入包括爆炸半径、可逆性和认知不确定性的三维风险分解,为操作员提供可解释的每次操作安全界面;三是设计上下文自适应的人在回路门,将升级从二元故障安全机制转变为响应值班负载和业务关键性的带宽感知控制层。完整策略从历史事件日志离线学习,实现对预期错误修复率的显式控制。在火车票微服务基准上的实验表明,相对于强大的运行手册基线,我们的框架将错误修复率降低了 39%,同时将修复成功率提高了 2.5 个百分点,并将值班升级负载降低了 17%。

英文摘要

In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all. Yet existing automated remediation systems are designed to generate actions rather than to decide whether intervention is warranted, leaving safety as an afterthought enforced by manual approval. This paper makes three contributions to close this gap: (i) we reformulate safe remediation as a risk-constrained intervention decision problem and cast it as a Constrained Markov Decision Process (CMDP), in which the agent maximizes repair success subject to a bounded false remediation rate (FRR); (ii) we introduce a three-dimensional risk decomposition comprising blast radius, reversibility, and epistemic uncertainty, providing operators with an interpretable per-action safety interface; and (iii) we design a context-adaptive human-in-the-loop (HITL) gate that turns escalation from a binary failsafe into a bandwidth-aware control layer responsive to on-call load and business criticality. The full policy is learned offline from historical incident logs, enabling explicit control of the expected FRR. Experiments on the Train Ticket microservice benchmark with Chaos Mesh fault injection and an RCAEval-aligned fault taxonomy show that our framework reduces FRR by 39% while improving repair success by 2.5 points over a strong runbook baseline, and reduces on-call escalation load by 17% relative to a fixed-threshold variant.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑