发表机构
Decentralized Science Lab, College of Computing and Software Engineering(分布式科学实验室,计算机与软件工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型生成代码中的拼凑问题,将结构一致性形式化并引入故障分类法,提出混合验证框架,经实证评估和真实世界验证,揭示多数结构故障难被现有检查发现,模型间故障模式有别,此类故障普遍存在。
AI 中文摘要
大语言模型生成的代码通常能编译、通过测试且看似正确,但部署后却会出错,根源往往是结构而非逻辑问题。每个补丁在局部有效但全局不一致,标准持续集成工具链很少暴露这些故障。本文将结构一致性形式化为存储库工件图形表示上的一致性不变量,引入八类故障分类法。提出混合验证框架,通过实证评估发现多数结构故障完全避开类型检查、测试和静态分析安全测试,且模型间故障模式有质的差异,真实世界验证证实这些故障普遍存在。
英文摘要
LLM-generated code often compiles, passes tests, and appears correct, yet breaks once deployed. The root cause is frequently structural rather than logical. A generated endpoint references configuration keys never declared in the project, an import targets a package that does not exist in any registry, or a new route omits the authentication guard applied to every sibling endpoint. Each patch is locally valid but globally incoherent, and standard CI toolchains rarely surface these failures. As LLM-powered coding tools see widespread adoption, this blind spot poses a growing risk to software quality. We call this the \textbf{patchwork problem}. This paper formalizes structural coherence as consistency invariants over graph representations of repository artifacts, including import, call, dependency, configuration, schema, resource, control-flow, and routing graphs, and introduces an eight-category failure taxonomy distinguishing defects specific to LLM generation from those merely amplified by it. We present a hybrid verification framework that delegates to mature static analysis tools where they already excel and deploys purpose-built detectors for cross-cutting invariants underserved by existing toolchains, targeting provable constraint violations rather than heuristic pattern matching. Empirical evaluation across two frontier models under four prompting strategies reveals that the vast majority of structural failures evade type checking, testing, and SAST entirely, and that failure patterns diverge qualitatively between models in ways that challenge model-agnostic mitigation strategies. External validation on real-world AI-generated repositories confirms that these failures are not artifacts of controlled experimentation but are prevalent wherever LLMs write code with minimal human oversight.