发表机构
Databricks(Databricks公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Consort是一个规格优先、测试驱动的智能体框架,通过确定性编排器和不可变测试等不可绕过的控制,在实时数据库分支上强制工程纪律,确保代码诚实可验证且可维护。
AI 中文摘要
当智能体编写代码时,开发框架便成为非确定性工作者的控制系统。自2025年以来,规格优先、智能体驱动的框架迅速获得关注;可安装的框架包括GitHub Spec Kit、obra/superpowers、BMAD和GSD,以及我们自己的框架,均通过规格或持久规划工件捕获意图。由于它们在预先捕获意图上达成一致,区分它们的关键在于各自如何强制实施工程纪律,以保持智能体编写的代码干净、正确且可维护。每个框架都以某种方式强制实施该纪律;差异在于方式。我们将其特征化为三种模式:通过说服强制(模型可能忽略的提示纪律)、通过前置结构强制(强规格,随后是可信构建),以及通过智能体无法编辑的控制强制(确定性编排器、人工批准的关卡、不可变测试,以及必须针对实时分支数据库通过的绿色结果)。我们引入Consort,一个基于第三种模式的规格优先、测试驱动智能体框架,通过智能体在其内部运行但无法绕过的控制来强制实施该纪律,其中确定性编排器驱动独立角色智能体,在实时数据库分支上通过规格优先设计通道和测试驱动构建通道。我们认为,在代码中强制实施测试和关卡可保持智能体编写的代码诚实且可验证,而其专门角色(如同之前的人类角色)则使其可维护,这些主张我们将其表述为预注册、可测试的假设。
英文摘要
When an agent writes code, the development framework becomes the control system for a non-deterministic worker. Spec-first, agent-driven frameworks have gained rapid traction since 2025; the installable ones, GitHub Spec Kit, obra/superpowers, BMAD, and GSD, and our own, all capture intent through a specification or durable planning artifacts. Since they agree on capturing intent up front, what separates them is how each enforces the engineering discipline that keeps agent-written code clean, correct, and maintainable. Every framework enforces that discipline somehow; they differ in how. We characterize three modes: enforcement by persuasion (prompt discipline the model may ignore), by front-loaded structure (strong specs, then a trusted build), and through controls the agent cannot edit (a deterministic orchestrator, human-approved gates, immutable tests, and a green result that must pass against a live, branched database). We introduce Consort, a spec-first, test-driven agent framework built on the third, enforcing that discipline through controls the agent runs inside but cannot bypass, in which a deterministic orchestrator drives separate role agents through a spec-first design lane and a test-driven build lane on a live database branch. We argue that enforcing the tests and gates in code keeps agent-written code honest and verifiable, while its specialized roles, like the human roles before them, are what make it maintainable, claims we frame as a pre-registered, testable hypothesis.
Comments9 pages, 2 figures, 3 tables, for associated framework, see https://github.com/databricks-solutions/consort