发表机构
Siemens Digital Industries Software(西门子数字工业软件)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对比5种架构规范格式对6种LLM编码智能体的代码生成质量影响,发现结构化格式可作为能力均衡器,显著提升弱模型性能,对成本优化部署价值最大。
AI 中文摘要
基于大语言模型(LLM)的编码智能体可根据高层描述生成完整软件系统,但架构规范的格式如何影响生成代码的质量,以及该影响是否取决于模型能力,目前尚不明确。我们开展了一项受控实验,对比5种信息等价的规范格式(非正式散文、带约束和架构决策记录(ADR)的Mermaid图、OpenAPI、C4/Structurizr领域特定语言(DSL)、带ArchUnit风格规则的TypeScript接口契约),涉及来自3个厂商系列的6种模型(Anthropic Claude、OpenAI GPT、Google Gemini)。在90次多轮智能体试验中,规范格式呈现出显著的“格式×模型”交互效应:最强模型(Sonnet 4.6、GPT-5)的格式影响极小,质量差幅为0.17-0.92;较弱模型的格式差幅达0.83-2.42分,其中贴近代码的格式(OpenAPI、TypeScript契约)可弥补大部分能力差距。中端模型若陷入较强模型可避免的编译调试循环,会比前沿模型消耗更多token却产出更差结果。自验证率在能力谱上从100%(Sonnet)降至0%(Gemini Flash)。TypeScript契约使最弱模型的API路由覆盖率从33%提升至100%,增幅达两倍。结构化架构规范可作为能力均衡器,其价值与模型强度成反比,在成本优化部署中收益最大。
英文摘要
LLM-based coding agents generate complete software systems from high-level descriptions, yet little is known about how the format of architecture specifications affects the quality of generated code or whether this effect depends on model capability. We present a controlled experiment comparing five informationally equivalent specification formats (informal prose, Mermaid diagrams with constraints and ADRs, OpenAPI, C4/Structurizr DSL, and TypeScript interface contracts with ArchUnit-style rules) across six models from three vendor families (Anthropic Claude, OpenAI GPT, Google Gemini). Across 90 multi-turn agent trials, specification format shows a strong format x model interaction. On the strongest models (Sonnet 4.6, GPT-5), format barely matters (quality spread 0.17-0.92). On weaker models, format produces spreads of 0.83-2.42 points, with code-proximate formats (OpenAPI, TypeScript contracts) recovering most of the capability gap. Mid-tier models can consume more tokens than frontier models for worse output when they enter compilation debugging loops that stronger models avoid. Self-validation rates collapse from 100% (Sonnet) to 0% (Gemini Flash) across the capability spectrum. TypeScript contracts triple API route coverage for the weakest model (33% to 100%). Structured architecture specifications serve as a capability equalizer, with value inversely proportional to model strength and the largest returns for cost-optimized deployments.