arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TriFleetRCA:面向 Kubernetes 的本地大语言模型根因分析

TriFleetRCA: On-Premise LLM Root Cause Analysis for Kubernetes

Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary

arXiv 2609.23766首次发表:更新:

发表机构

Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出 TriFleetRCA,一个在单块本地 GPU 上运行的 Kubernetes 根因分析流水线,通过三范围证据收集、模板去重与 BM25 排序及运维手册防护,在实时集群注入故障下达到 0.85-0.95 命中率,并实现防御纵深。

AI 中文摘要

远程站点的根因分析速度缓慢:证据分散在 Pod 日志、Kubernetes 事件和集群级对象中,而且许多运维人员根本无法将生产日志发送给托管模型。本地推理消除了第二个限制,但提出了一个实时集群基准测试尚未解决的问题:当一台工作站 GPU 同时解决了模型和上下文预算问题时,应如何检索证据?当模型查阅的运维手册(runbook)被篡改时,会发生什么?我们提出了 TriFleetRCA,一个完全在单块本地 GPU 上运行的流水线,它从三个范围(Pod、命名空间、集群)之一收集证据,通过模板去重后按 BM25 排序,通过摄取防护(ingest guard)过滤运维手册,并返回根因及其支持的证据行。我们在一个实时 Kubernetes 集群上进行了评估,向其中注入了四个故障,因此真实结果是通过构造已知的,使用 Qwen2.5-14B-Instruct 在温度为 0 的情况下进行了 100 次分析。在 Pod、命名空间和集群范围下的命中率分别为 0.85、0.90 和 0.95;区间有重叠,但整体范围效应来自一个故障,其原因是集群级对象,而集群范围多消耗了 55% 的令牌。排序前的去重将命中率从 0.75 提高到 0.90,且令牌成本相同。一个指示模型删除命名空间的恶意运维手册每次运行都被防护拒绝;在禁用防护的情况下,模型在所有 20 次分析中都拒绝遵循该指令,这使得防护成为纵深防御而非唯一屏障。将引用质量与准确性分开被证明是有信息量的:一个故障在每次试验中都被正确诊断但引用错误,这是一种准确性掩盖的失败模式。中位延迟为 1.6 秒,对应 2,200 个提示令牌。我们发布了该流水线、故障注入器及所有记录。

英文摘要

Root cause analysis at a remote site is slow: evidence is scattered across pod logs, Kubernetes events and cluster-level objects, and many operators cannot send production logs to a hosted model at all. On-premise inference removes the second constraint but raises a question live-cluster benchmarks have not addressed: when one workstation GPU fixes both the model and the context budget, how should evidence be retrieved, and what happens when the runbooks the model consults have been tampered with? We present TriFleetRCA, a pipeline running entirely on one on-premise GPU that collects evidence at one of three scopes (pod, namespace, cluster), ranks it by template de-duplication then BM25, filters runbooks through an ingest guard, and returns a root cause with the evidence lines supporting it. We evaluate on a live Kubernetes cluster into which we inject four faults, so ground truth is known by construction, across 100 analyses with Qwen2.5-14B-Instruct at temperature 0. The hit rate was 0.85, 0.90 and 0.95 at pod, namespace and cluster scope; intervals overlap, but the whole scope effect comes from the one fault whose cause is a cluster-level object, and cluster scope costs 55% more tokens. De-duplication before ranking raised the hit rate from 0.75 to 0.90 at equal token cost. A poisoned runbook telling the model to delete the namespace was rejected by the guard every run; with the guard disabled the model declined to follow it in all 20 analyses, making the guard defence in depth rather than the sole barrier. Separating citation quality from accuracy proved informative: one fault was diagnosed correctly and cited incorrectly every trial, a failure mode accuracy conceals. Median latency was 1.6 s at 2,200 prompt tokens. We release the pipeline, the fault injector and all records.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑