arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI辅助软件开发的治理方法论层:缺陷分类、受控 ablation 及流程优先于能力的证据

A Governance Methodology Layer for AI-Assisted Software Development: Defect Taxonomy, Controlled Ablation, and a Test of Process-Over-Capability

Sungjin Kwon

arXiv 2609.04218首次发表:更新:

AI 中文总结

该研究针对AI辅助软件开发的语义缺陷检测缺口,提出含缺陷分类、治理门、方法论代码化的治理层,经受控 ablation 实验证实审查流程设计比模型能力更影响关键缺陷覆盖率。

AI 中文摘要

自主编码智能体生成的输出能快速通过编译、类型安全、CI等语法检查,但语法正确不代表语义正确,设计边界、安全不变量和可维护性约定仍是自动化流程无法检测的结构。本文为缩小该差距作出四项贡献:第一,提出基于五个AI智能体权限与治理模块的缺陷分类法,区分静态分析可结构检测的缺陷与需语义审查的缺陷;第二,描述运行时解耦的治理门——一种基于文件的协议,读取生成器输出并输出结构化判定结果,无需API耦合,可跨代码生成工具移植;第三,将方法论形式化为代码:把验证协议表达为可版本控制、可执行的跨平台产物,采用两层架构分离可移植方法论与主机特定自动化;第四,开展受控 ablation 实验(E-ablation,N=5个制品,8项独立真实值),对比 harness 结构化审查与同制品上 token 匹配的非结构化审查提示。结构化条件对宽松真实值的召回率为62%,非结构化条件为50%;严格召回率为25%,非结构化条件为0%,且结构化条件给出的严重程度等级更一致,非结构化条件会夸大严重程度。两种条件均遗漏人工QA审查员识别的文档质量缺陷,表明结构化AI审查与人工流程检查具有互补性。在该样本范围内,结果支持以下论点:对于AI辅助软件开发,审查流程设计而非审查者模型能力是关键缺陷覆盖率的主导因素。

英文摘要

Autonomous coding agents produce output that passes syntactic checks -- compilation, type safety, CI -- at high velocity. Yet syntactic correctness does not imply semantic correctness: design boundaries, security invariants, and maintainability contracts remain structurally invisible to automated pipelines. This paper makes four contributions. First, we present a defect-class taxonomy grounded in five AI agent permission and governance modules, separating defects structurally detectable by static analysis from those requiring semantic review. Second, we describe a runtime-decoupled governance gate -- a file-based protocol that reads generator output and emits a structured verdict without API coupling, hence portable across generators. Third, we formalize methodology-as-code: expressing a verification protocol as a version-controlled, executable artifact whose two layers separate portable methodology from host automation. Fourth, we report a controlled ablation experiment (E-ablation, N=5 artifacts, 8-item independent ground truth) comparing harness-structured review against a token-matched unstructured prompt. The structured condition records 62% lenient recall against 50%, with 25% strict against 0%. A severity-grade differential reported earlier does not survive blind re-grading and is withdrawn (Sec. 6.6). An independent-session re-test with blind scoring does not replicate that contrast: the conditions differ by one strict hit in 24 (6/24 against 5/24), the structured aggregate again 25% and the unstructured 0% not recurring (Sec. 6.7). Both conditions miss document-quality defects identified by a human QA reviewer, indicating complementarity between structured AI review and human process inspection. The re-test does not distinguish structured review from a detailed unstructured prompt here, so process design as the dominant factor remains a hypothesis, not a result of this paper.

Commentsv3: corrects three body passages the v1.2.2 downgrade did not reach. Sec. 1.2 C4 is restated as a test whose contrast does not survive Sec. 6.7, not as evidence. The close of Sec. 6.6 no longer says no replication has tested the recall result, since Sec. 6.7 has. Sec. 11 no longer claims all artifacts are released, matching Sec. 6.4 and Apps. A-B. No results or numbers changed. 23 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑