采用AI编码智能体的规范优先收敛:在71.7万行代码库中拆解189个文件的核心架构不变量的案例研究
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
浏览论文内容
中文总结 AI 辅助
本文通过案例研究,展示AI编码智能体在规范优先协议下,成功在71.7万行代码库中拆解核心架构不变量,耗时3天、成本2430美元,修正201个缺陷,涉及189个文件。
中文摘要 AI 辅助
本文报告了一项完全可观测的大规模架构重构案例研究,该研究由AI编码智能体在规范优先协议下完成,生成代码无人工审核,且无预存的测试预言机(test oracle)来验证目标行为。该任务是拆解大型相互依赖代码库中的核心不变量,作者评估通过增量重构几乎无法完成,这类变更通常需要重写代码。在本文所述协议下,智能体成功完成了任务。该系统是一个包含3648个文件、717725行代码的生产级TypeScript应用。任务要求拆解核心生命周期不变量:即UI面板在AI请求期间保持打开状态的保证。目标行为是流式生成过程在其面板关闭后仍能继续,重新打开时可重新附加到同一活跃流,且无丢失或重复。协议流程为:智能体制定形式化规范,进行14次精化循环,将该规范与源代码进行审计;随后执行原子实现,再通过编译/测试反馈循环,接着进行17次验证循环,将代码与冻结的规范进行审计。在全部31次审计过程中,智能体在人工运行程序前已修正了201个缺陷。收敛准则为经验性的:连续两次验证过程返回零缺陷。该变更涉及189个文件(其中31个为新文件),加上提取阶段,两次提交共涉及288个文件,新增34770行代码,删除16422行代码。在首次及后续约30次会话中,软件均符合规范,未观测到任何bug。耗时3天,成本2430美元。完整规范和原始会话日志(1500多页法语内容)已作为证据发布,可供检查流程并提交给语言模型进行一致性检查。
英文摘要
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request. The target behaviour was that a streaming generation survives the closing of its panel and can be reattached, on reopening, to the same live stream with no loss or duplication. The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program. The convergence criterion was empirical: two consecutive verification passes returning zero findings. The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430. The full specification and raw session logs, 1,500+ pages in French, are published as evidence, allowing inspection of the process and submission to a language model for consistency checking.
发表机构
- AI Sovereign Labs(AI主权实验室)
机构由 AI 辅助整理,请以论文原文为准。