arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22406cs.SE

Vibe编码:一项测试驱动开发的实验

Vibe Coding: An Experiment with Test-Driven Development

Moritz Mock, Barbara Russo

首次发表
浏览论文内容

中文总结 AI 辅助

研究探索人类与CLLMs通过vibe编码对等协作,设计四个交互模型并实现相应TDD工作流程。通过实验发现交互模型选择取决于开发目标,智能工作流程利于快速开发但有复杂性,协作工作流程产生更高质量测试套件。

中文摘要 AI 辅助

背景:对话式大语言模型(CLLMs)可通过与用户自然语言协作自动生成代码,但协作不佳会导致输出质量差。目的:本探索性研究旨在调查人类和CLLMs如何通过vibe编码作为对等方进行协作,该方法整合了提示工程、敏捷设计和人机共创原则以增强协作。方法:设计了四个代表软件开发过程中不同协作模式的交互模型,基于这些模型用结构化提示和Python脚本实现相应的测试驱动开发(TDD)工作流程。对TDD专业人员进行了对照预实验研究以比较单独和协作工作流程,还对相同开发任务进行了全自动和智能工作流程的重复探索性执行以获取补充证据。结果:研究结果表明交互模型的选择应取决于开发目标。智能工作流程最适合快速开发和功能正确的生产代码,但可能会引入额外的实现复杂性,也可能引入功能规范未明确要求的额外实现决策,导致未测试的决策点。相比之下,协作工作流程产生更高质量、组织更好的测试套件。结论:我们的工作探索了人类和CLLMs如何通过vibe编码作为对等方进行协作,交互模型的选择应取决于开发目标。

英文摘要

Context: Conversational Large Language Models (CLLMs) can automatically generate code by collaborating with users through natural language. However, poor collaboration can lead to poor quality output. Objective: This exploratory study aims to investigate how humans and CLLMs can collaborate as peers through vibe coding, an approach that integrates principles from prompt engineering, agile design, and human-AI co-creation to enhance collaboration. Method: We designed four interaction models representing different collaboration patterns in the software development process: the solo model (human-only development), the collaborative model (human-CLLM collaboration), the fully automated model (development autonomously performed by a CLLM), and the agentic model (development autonomously performed by the MetaGPT~X platform). Based on these models, we implemented corresponding Test-Driven Development (TDD) workflows using structured prompts and Python scripts. We then conducted a controlled pre-experimental study with TDD professionals to compare the solo and collaborative workflows. In addition, we performed repeated exploratory executions of fully automated and agentic workflows on the same development tasks to obtain complementary evidence. Results: Our findings suggest that the choice of interaction model should depend on the development objective. Agentic workflows are best suited for rapid development and functionally correct production code but may introduce additional implementation complexity. However, they may also introduce additional implementation decisions that are not explicitly required by the functional specifications, resulting in untested decision points. In contrast, collaborative workflows produce higher-quality, better-organized test suites. Conclusions: Our work explored how...

↑