arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

架起行为与实现的桥梁:用于行为驱动开发的自动化Java胶水代码生成

Bridging Behavior and Implementation: Automated Java Glue Code Generation for Behavior-Driven Development

Xinyu Shi, Zhou Yang, An Ran Chen

arXiv 2607.19703首次发表:更新:

AI 中文总结

研究针对行为驱动开发中胶水代码开发维护难题,提出分层多智能体框架AutoGlue,遵循行为优先工作流程。经实验评估,相比少样本提示,其在代码生成指标上有显著提升,能有效连接行为规范与项目代码,支持软件开发。

AI 中文摘要

行为驱动开发(BDD)通过自然语言场景帮助技术和非技术利益相关者对软件需求达成共识。胶水代码通过将每个步骤映射到相应的项目代码使这些场景可执行。然而,开发和维护胶水代码需要行为和底层代码库的知识,随着需求演变,这成为BDD中劳动密集的部分。虽然大语言模型(LLMs)已显示出强大的代码生成能力,但用于自动化胶水代码生成尚未被探索。我们提出AutoGlue,一个用于自动化Java胶水代码生成的分层多智能体框架。AutoGlue遵循行为优先的工作流程,分离行为解释、上下文检索和代码生成。我们在来自八个开源Java项目的1307个步骤上评估AutoGlue。与少样本提示相比,AutoGlue将API F1提高了58.7%,将CodeBLEU提高了43.7%。它为46.1%的评估步骤生成直接可用的胶水代码,大多数部分正确的输出只需少量修订。消融结果表明行为解释和项目感知上下文检索对生成质量都有很大贡献。这些发现表明LLMs可以有效地将自然语言行为规范与项目代码连接起来,并支持规范驱动的软件开发。

英文摘要

Behavior-Driven Development (BDD) helps technical and non-technical stakeholders share a common understanding of software requirements through natural-language scenarios. Glue code makes these scenarios executable by mapping each step to the corresponding project code. However, developing and maintaining glue code requires knowledge of both the intended behavior and the underlying codebase, making it a labor-intensive part of BDD as requirements evolve. Although large language models (LLMs) have shown strong code generation capabilities, their use for automated glue code generation remains unexplored. This task requires reasoning over underspecified behavior, related BDD artifacts, and large project codebases. We present AutoGlue, a hierarchical multi-agent framework for automated Java glue code generation. AutoGlue follows a behavior-first workflow that separates behavior interpretation, context retrieval, and code generation. A Behavior Interpreter derives the intent of a step from its scenario context, while a Developer agent retrieves relevant BDD artifacts and project code before generating the final glue code. We evaluate AutoGlue on 1,307 steps from eight open-source Java projects. Compared with few-shot prompting, AutoGlue improves API F1 by 58.7% and CodeBLEU by 43.7%. It produces directly usable glue code for 46.1% of the evaluated steps, while most partially correct outputs require only minor revisions, such as adding missing actions or refining parameters. Ablation results show that behavior interpretation and project-aware context retrieval both contribute substantially to generation quality. These findings demonstrate that LLMs can effectively connect natural-language behavior specifications with project code and support specification-driven software development.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑