发表机构
Singapore Management University; University College London; Tencent; University of Alberta; Alberta Machine Intelligence Institute(新加坡管理大学; 伦敦大学学院; 腾讯; 阿尔伯塔大学; 阿尔伯塔机器智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过混合方法研究从业者构建软件工程代理的方式,发现随着实现成本降低瓶颈转移,刻画了七阶段工作流程和评估驱动开发转变,还识别出团队面临的六个挑战及应对实践。
AI 中文摘要
软件工程(SE)代理的兴起,即基于大语言模型的代理,能够理解大型代码库并在有限的人工干预下执行工程任务,其发展迅速且应用广泛。但对于开发者在实践中如何构建这些系统却知之甚少:现有研究多挖掘仓库或考察部署,很少研究SE代理的构建方式。本文通过对12个组织的20名从业者进行半结构化访谈以及对80名从业者进行在线调查,首次研究了在SE代理开发中SE流程如何变化以及开发者面临哪些挑战。我们发现随着实现成本降低,瓶颈发生转移而非消失:诸如需求、协调、审查和部署等长期存在的非编码工作变得更加突出,而审查和评估代理输出成为新的核心工作。我们刻画了一个七阶段工作流程以及向评估驱动开发的转变,其中评估引导迭代,规范成为人类和代理都能读取的版本化工件。我们还进一步识别了团队面临的六个挑战以及他们为应对这些挑战所采用的实践,包括不可靠的评估信号、代码增长超过理解速度导致的理解债务以及供应商端模型更新带来的行为变化。
英文摘要
The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little is known about how developers build these systems in practice: existing studies mine repositories or examine deployment, but few investigate how SE agents are constructed. Through semi-structured interviews with 20 practitioners from 12 organizations and an online survey of 80 practitioners, this paper is the first to study how SE processes are changing in the development of SE agents and what challenges developers face. We find that as implementation becomes cheaper, bottlenecks shift rather than disappear: long-standing work in requirements, coordination, and deployment becomes more visible, while reviewing generated code and evaluating agent behavior become new and increasingly central forms of work. We characterize a seven-stage workflow and five process shifts, including a move toward evaluation-driven development, in which evaluation is increasingly defined early and steers iteration, and the emergence of specifications as first-class artifacts that teams test and version alongside code. We further identify six challenges that teams face, together with 12 corresponding practices they use or propose to address them, including unreliable evaluation signals, comprehension debt as code outpaces understanding, and behavioral changes introduced by provider-side model updates.