发表机构
Teddysoft; National Taipei University of Technology(泰迪软件; 国立台北科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM生成代码不满足需求的问题,基于亚历山大形式理论提出不匹配治理开发(MGD)方法论,通过显式表示、检测器、修复循环和人类门控回路,实现LLM代码的可靠保障。
AI 中文摘要
由大型语言模型(LLM)生成的代码不能假定其满足既定需求。审查、测试和静态分析仍然适用,但一个充分的保障机制需要其中哪些手段,以及它们各自扮演何种角色,仍是开放问题。我们提出了一种源自克里斯托弗·亚历山大形式理论的设计理论,以及一套应用该理论的方法论。在亚历山大的论述中,形式与其情境之间的适配只能通过识别出的不匹配项的缺失来消极感知。我们使组织的传统变得明确,并从中以及从问题的分类中推导出不匹配项。该理论将LLM建模为一种非本土的乡土建造者,它训练于众多代码库但对任何一个都不熟悉,其输出倾向于偏向主流惯例而非本地传统。我们设计了四种机制:问题的显式表示(杰克逊的问题框架)和传统的显式表示(四种形式的模式语言);确定性的不匹配检测器;修复循环;以及一个由人类把关的立法回路,用于管理表示和检测器。我们将由此产生的方法论——一种保障工程实践——称为不匹配治理开发(MGD)。其双循环过程将自主内循环(LLM在此迭代对抗门控)与人类外循环(在此根据现实世界评判规格)分离开来。它们共同构成S=P=T=W保障模型(规格、程序、测试、世界),其中的等号表示关系而非同一性。我们报告了从64个问题框架规格构建和重建一个由四个事件溯源聚合组成的Scrum系统的证据,该系统由约1300个生成的测试和28个阻塞门控验证,其中一个门控应用了188条规则。这回应了亚历山大1996年OOPSLA挑战的生成性维度。道德维度——即规格是否仍然适配现实世界——需要人类判断,属于外循环。
英文摘要
Code generated by large language models (LLMs) cannot be assumed to meet specified requirements. Reviews, testing, and static analysis still apply, but which of them a sufficient harness needs, and in what role, is open. We propose a design theory derived from Christopher Alexander's theory of form, and a methodology for applying it. In Alexander's account, fit between a form and its context can be perceived only negatively, through the absence of identified misfits. We make the organization's tradition explicit and derive the misfits from it and from the problem's classification. The theory models the LLM as a non-native vernacular builder, trained on many codebases but native to none, whose output tends to drift toward mainstream conventions rather than the local tradition. We engineer four pieces of machinery: explicit representations of the problem (Jackson's problem frames) and of the tradition (a four-form pattern language); deterministic misfit detectors; a fix loop; and a human-gated legislative circuit governing the representations and detectors. We call the resulting methodology, a practice of harness engineering, Misfit-Governed Development (MGD). Its dual-loop process separates an autonomous inner loop, where the LLM iterates against the gates, from a human outer loop, where specifications are judged against the world. Together they form the S = P = T = W assurance model (specification, program, tests, world), whose equals signs name relations, not identity. We report evidence from building and rebuilding a Scrum system of four event-sourced aggregates from 64 problem-frame specifications, verified by about 1,300 generated tests and 28 blocking gates, one applying 188 rules. This addresses the generativity dimension of Alexander's 1996 OOPSLA challenge. The moral dimension, whether the specification still fits the world, requires human judgment and belongs to the outer loop.