arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

比较 vibe 编码工具生成代码的质量

Comparing the Quality of Code Generated by Vibe Coding Tools

Gustavo da Mota, Kiev Gama

arXiv 2608.16302首次发表:更新:

AI 中文总结

本研究对比 Lovable、v0、Replit 三种 vibe 编码工具生成代码的结构质量,发现三者存在不同质量特征,选择工具需考量结构层面的权衡。

AI 中文摘要

在软件开发中,使用智能体(agent)进行自动代码生成已越来越普遍,但生成代码的质量仍存在诸多担忧,包括可维护性、可读性及长期演进等方面。本研究比较了三种广泛使用的 vibe 编码工具——Lovable、v0 和 Replit 从单一生成提示词出发所生成代码的结构质量。我们为每种工具生成三个独立项目,共九个 Web 应用程序,并将其提交给 SonarQube 进行静态分析。我们收集的指标包括问题数量、严重程度分布、预计修复工作量、圈复杂度和认知复杂度以及代码重复度。初步结果显示,这些工具呈现出不同的定性特征:Lovable 的问题集中在较低严重程度,但每千行代码(KLOC)的代码异味密度明显更高;而 v0 和 Replit 生成的代码具有更严重的严重程度分布。这些发现表明,选择 vibe 编码工具涉及超出感知生产力的结构权衡。

英文摘要

The use of AI agents for automatic code generation has become increasingly common in software development. However, concerns remain about the quality of the generated code, including aspects of maintainability, readability, and long-term evolution. This study compares the structural quality of code produced by three widely adopted vibe coding tools --- Lovable, v0, and Replit --- starting from a single generation prompt. We generate three independent projects per tool, totalling nine web applications, and submit them to static analysis with SonarQube. We collect metrics such as the number of issues, severity distribution, estimated remediation effort, cyclomatic and cognitive complexity, and code duplication. Preliminary results show that the tools exhibit distinct qualitative profiles: Lovable concentrates issues of lower severity but presents a substantially higher density of code smells per KLOC, while v0 and Replit produce more code with more aggressive severity profiles. These findings suggest that choosing between vibe coding tools involves structural trade-offs that go beyond perceived productivity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑