发表机构
Hochschule Offenburg(奥芬堡应用技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文实现了规范增长引擎,通过两层结构防止规范-代码分歧,并借助由模型扮演的三类智能体扩展规范,设置三类开关形成18种运行模式,核心主张为自主运行的价值取决于确定性实例的衡量。
AI 中文摘要
规范增长引擎将AI辅助软件开发锚定在与代码耦合的规范图中。本文描述其实现,该实现服务于两项任务并将其划分为两层。第一层防止规范-代码分歧:确定性引擎验证规范图,将其与代码的导入图进行比较,从记录的测试证据中获取节点的验证状态,并根据变更可能造成的破坏对每项变更进行分类,全程不调用模型。第二层通过智能体扩展规范:意图作者、规划器和编码器分别由各自的模型扮演,它们循环扩展图,且每轮后由确定性规则决定是否继续运行。人类判断的委托程度由三个独立开关设置——草稿门、破坏性变更委托和运行模式,结合两种项目基础搭建方式,共产生18种项目运行方式。我们描述了每种方式、智能体与人类通信的门、请求和豁免,以及从完全手动工作到无监督运行(人类事后审查其决策)的操作范围。整个设计遵循一个核心主张:自主运行的价值仅等同于衡量它的确定性实例。
英文摘要
The Spec Growth Engine anchors AI-assisted software development in a graph of specifications that the code is coupled to. This paper describes its implementation, which serves two tasks and keeps them apart as two layers. The first layer prevents spec-code divergence: a deterministic engine validates the spec graph, compares it with the code's import graph, earns a node's verified status from recorded test evidence, and classifies every change by what it can break -- without calling a model. The second layer grows the spec with agents: an intent author, a planner and a coder, each played by its own model, extend the graph in rounds, and a deterministic rule decides after each round whether the run goes on. How much of the human's judgement is delegated is set by three independent switches -- a draft gate, a delegation for breaking changes, and the run mode -- which, with two ways of laying a project's floor, give eighteen ways to run a project. We describe each of them, the gates, requests and waivers through which agents and the human communicate, and the spectrum of operation from entirely manual work to an unsupervised run whose decisions the human reviews afterwards. Throughout, one claim holds the design together: an autonomous run is worth only as much as the deterministic instance that measures it.
Comments16 pages, 3 figures, 8 tables. Code: https://github.com/grabow/spec-growth-engine