发表机构
Independent Researcher
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对GitOps中LLM修复方案的安全问题,提出将LLM字段变更意图与确定性YAML编辑结合的方案,实现安全高效的GitOps修复并发布基准。
AI 中文摘要
LLM智能体越来越多地用于诊断故障并提出修复方案。在GitOps工作流中,应用修复意味着编辑受版本控制的配置文件,最直观的实现方式是让模型生成编辑后的文件或差异,这也是从业者最先想到的方案。通过对真实Kubernetes清单评估该方案,我们发现没有任何文本生成策略可用于无人值守自动化。统一差异存在安全问题:在严格补丁下几乎无法应用,但这是假象,因为兼容工具GNU patch可应用96%的差异,却会无声地错误应用约1/7(14%-20%)的差异且无错误信号。全文件重写依赖模型能力:小型模型会损坏文件,而前沿模型通常正确但具有不确定性(部分运行中会无声地删除字段或编辑相邻内容),必须重新生成整个文件,每次编辑的成本为O文件大小。我们提出一种替代方案,将语义决策(资源、字段和值)与文件编辑的语法操作分离。智能体仅输出结构化字段变更意图;确定性管道通过YAML解析器的节点位置标记,按(类型、名称)索引清单,定位目标标量的精确字符跨度,并仅替换原始文本中的该跨度。由于文件从未重新序列化,差异天生最小,格式和注释得以保留,编辑正确且与模型无关的确定性,生成成本为O(1)。我们的贡献是将LLM提出的意图与GitOps的确定性、故障关闭应用契约相结合,在KubeAstra(Apache-2.0许可证)中实现该方案并发布基准。我们的主张范围是忠实地应用已知变更;变更是否正确由人类PR审查决定。
英文摘要
LLM agents increasingly diagnose incidents and propose remediations. In a GitOps workflow, applying a fix means editing a version-controlled config file, and the obvious implementation, having the model author the edited file or a diff, is what practitioners reach for first. Evaluating that choice on real Kubernetes manifests, we find no text-generation strategy is safe for unattended automation. Unified diffs are unsafe: under strict patching almost none apply, but that is an artifact, since a tolerant tool (GNU patch) applies 96%, yet silently misapplies about 1 in 7 (14-20%) with no error signal. Full-file rewrite is capability-dependent: a small model corrupts the file, while a frontier model is usually correct but non-deterministic (it silently drops a field or edits a neighbor on some runs) and must regenerate the whole file, costing O(file size) per edit. We present an alternative that separates the semantic decision (which resource, field, and value) from the syntactic act of editing the file. The agent emits only a structured field-change intent; a deterministic pipeline indexes manifests by (kind, name), locates the target scalar's exact character span via the YAML parser's node position marks, and replaces only that span in the raw text. Because the file is never re-serialized, the diff is minimal by construction, formatting and comments are preserved, and the edit is correct and deterministic independent of the model, at O(1) generation cost. The contribution is the pairing of an LLM-proposed intent with a deterministic, fail-closed application contract for GitOps. We implement it in KubeAstra (Apache-2.0) and release the benchmark. Our claim is scoped to faithful application of a known change; whether the change is right is left to human PR review.
Comments8 pages, 2 figures, 5 tables, 3 appendices. Benchmark artifact: github.com/astraverse-io/kubeastra-bench (Apache-2.0). Implementation: github.com/astraverse-io/KubeAstra