Vibe编程:实践、性能、生产力与风险——一项最新进展综述
Vibe Coding: Practice, Performance, Productivity, and Risk -A State-of-the-Art Review
AI总结:
本综述针对2025年提出的Vibe编程,梳理跨学科实证证据,发现其代码生成可靠但故障检测弱,生产力结果因测量方式等差异存在分歧,还存在安全、版权及技能萎缩问题,提出新代码收益真实、成熟代码收益收缩的猜想。
AI中文摘要:
Vibe编程是一种AI辅助软件开发模式,开发者以自然语言描述开发意图,通过运行生成的代码而非阅读代码来验证结果,该概念由Andrej Karpathy于2025年2月提出,在17个月内就产生了首批实证证据。本最新进展综述汇集了软件工程、人机交互、劳动经济学、安全研究、治理及教育等跨学科领域的相关证据,对模型格局、工具生态系统及按任务类型划分的性能记录进行了调研,发现早期基准已达饱和状态,但任务级能力参差不齐:代码生成可靠,但故障检测能力弱,且代码文档难以审计。生产力记录起初相互矛盾:同行评审的现场实验显示每周任务量增加26%,独立随机试验则显示速度放缓19%,团队级遥测数据显示代码审查时间增加441%。我们认为,若保持测量方法、范围和时间跨度一致,这些结果是一致的,并确定了导致差异的六种模式,包括更广泛测量下的效应收缩、自我报告与独立测量的偏差、产出量与生产力的混淆,以及长期测试后大胆主张被撤回。我们还记录了已部署应用中的安全故障、大规模代码及开发者遥测数据中可见的代码质量下降、未解决的版权风险,以及技能萎缩的证据。本综述最后列出了未解决的研究问题和一项可证伪的猜想:在新代码上Vibe编程的收益是真实的,而在成熟代码库上收益会收缩或逆转,这将解释记录中的大部分分歧。
英文摘要:
Vibe coding - AI-assisted software development in which the developer describes intent in natural language and validates results by running rather than reading the generated code - was named by Andrej Karpathy in February 2025 and produced its first body of empirical evidence within seventeen months. This state-of-the-art review assembles that evidence across a cross-disciplinary corpus spanning software engineering, human-computer interaction, labour economics, security research, governance, and education. We survey the model landscape, the tool ecosystem, and the performance record by task type, finding the early benchmarks saturated but task-level capability uneven: reliable code generation alongside weak fault detection and hard-to-audit documentation. The productivity record is at first contradictory: peer-reviewed field experiments report +26% more tasks per week, independent randomised trials measure a 19% slowdown, and team-level telemetry shows code-review time up +441%. We argue these readings are consistent once measurement method, scope, and time horizon are held constant, and identify six patterns behind the dispersion, among them effect-shrinkage under broader measurement, self-report diverging from independent measurement, output volume conflated with productivity, and bold claims walked back once tested over longer horizons. We further document security failures in deployed applications, code-quality degradation visible in large-scale code and developer telemetry, unsettled copyright exposure, and evidence of skill atrophy. The review closes with the open research questions and one falsifiable conjecture: that the gains are real on new code and shrink or reverse on mature codebases, which would account for most of the disagreement in the record.