arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARGUS:基于MCP的Kubernetes事件根本原因分析工具

ARGUS: MCP-Grounded Root Cause Analysis for Kubernetes Incidents

Ergi Senja, Seyed Mohammad Reza Razavi Zadegan, Philipp Leitner

arXiv 2608.23084首次发表:更新:

AI 中文总结

ARGUS是一款基于MCP的Kubernetes事件根本原因分析助手,通过标准化MCP服务器关联LLM与实时可观测数据,在Slack频道提供诊断摘要,经评估其能100%识别根本原因,但修复建议可信度不足。

AI 中文摘要

Kubernetes事件分类需要关联来自指标、日志、容器状态以及跨多个监控工具的消息系统的信号,这一碎片化的工作流程会减慢诊断速度并加剧警报疲劳。大语言模型(LLM)在自动化根本原因分析(RCA)方面展现出潜力,但现有系统依赖定制的、特定于系统的数据访问层,无法在不同组织间复用。我们提出ARGUS,这是一款基于MCP的RCA助手,它通过标准化的MCP服务器将商用LLM与实时Kubernetes可观测性数据连接起来,这些服务器覆盖Kubernetes状态、Prometheus指标、Loki日志和NATS消息,并在值班工程师已在使用的Slack事件频道中提供结构化诊断摘要。我们采用三种互补方法对ARGUS进行初步评估:在10个Kubernetes事件场景中进行受控故障注入,从三个维度对生成的RCA摘要进行基于评分规则的评分,以及对某工业合作伙伴的6名值班工程师进行半结构化访谈。ARGUS在所有10个场景中都命名了正确的根本原因,综合MCP成功率为0.91。从业者信任诊断输出,但始终对推荐的修复措施表示怀疑。我们的核心发现是诊断/规定不对称性:ARGUS能可靠识别出什么出了问题,但在规定接下来该做什么方面被认为可靠性或可信度较低。这种模式在三种评估方法中均可见,对未来的自主智能体事件处理系统具有重要意义。

英文摘要

Kubernetes incident triage requires correlating signals from metrics, logs, container state, and messaging systems across multiple monitoring tools, a fragmented workflow that slows diagnosis and contributes to alert fatigue. Large language models (LLMs) have shown promise for automated root cause analysis (RCA), but existing systems rely on custom, system-specific data access layers that cannot be reused across organisations. We present ARGUS, an MCP-grounded RCA assistant that connects a commercial LLM to live Kubernetes observability data through standardised MCP servers covering Kubernetes state, Prometheus metrics, Loki logs, and NATS messaging, and delivers structured diagnostic summaries inside the Slack incident channel where on-call engineers already work. We conduct a preliminary evaluation of ARGUS using three complementary methods: controlled fault injection across ten Kubernetes incident scenarios, rubric-based scoring of the resulting RCA summaries on three dimensions, and semi-structured interviews with six on-call engineers at an industrial partner. ARGUS named the correct root cause in all ten scenarios with an aggregate MCP success ratio of 0.91. Practitioners trusted the diagnostic output but consistently expressed scepticism toward the recommended fixes. Our central finding is a diagnostic/prescriptive asymmetry: ARGUS reliably identifies what went wrong, but is perceived as less reliable or trustworthy at specifying what to do next. This pattern can be observed across all three evaluation methods, and has important implications for future autonomous agentic incident handling systems.

Comments11 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑