arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29208cs.SE

需求气味对基于LLM的代码生成的影响

On the Impact of Requirement Smells in LLM-Based Code Generation

Hugo Villamizar, Jannik Fischbach, Mert Şahin, Julian Frattini, Alessio Ferrari, Alexander Korn, Andreas Vogelsang, Daniel Mendez

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过扩展先前数据集,评估需求气味对LLM生成代码功能正确性的影响,发现气味密度增加降低正确性,且影响因任务而异,凸显需求质量的重要性。

中文摘要 AI 辅助

软件需求通常被纳入用于LLM辅助软件开发中的提示词。近期研究表明,需求气味会影响需求与代码之间的自动可追踪性,但关于其在代码生成中的影响的实证证据仍然有限。为弥补这一空白,我们基于先前一项关于自动可追踪性的研究,复用了其数据集和需求气味分类法,并扩展以评估LLM生成代码的功能正确性。利用一个包含四个应用的需求及相应系统测试的基准,我们逐步将语义、句法和词汇气味引入原本清晰的需求中,并分析其对生成实现的影响。结果表明,气味密度的增加通常与基于测试套件的功能正确性降低相关,尽管无气味的需求仍可能产生有缺陷的代码。我们还发现,不同气味类别的影响相似。这些发现为需求质量在LLM辅助代码生成中的重要性提供了额外的实证证据,同时表明高质量需求本身并不能保证正确性,因为正确性还取决于包括LLM在内的多个因素。与先前工作相比,我们的结果表明,需求气味的影响取决于软件工程任务:虽然其对可追踪性的影响不大,但代码生成似乎更为敏感。总体而言,这项工作激发了对LLM辅助软件工程中任务依赖性质量效应的进一步研究。

英文摘要

Software requirements are typically incorporated into prompts used in LLM-assisted software development. Recent work has shown that requirement smells can affect automated traceability between requirements and code, but empirical evidence on their effects in code generation remains limited. To address this gap, we build upon a prior study on automated traceability by reusing its dataset and requirement smell taxonomy, while extending it to evaluate the functional correctness of LLM-generated code. Using a benchmark consisting of requirements and corresponding system tests for four applications, we progressively introduced semantic, syntactic, and lexical smells into otherwise clear requirements and analyzed their influence on generated implementations. Our results suggest that increasing \textit{smell density} was generally associated with lower test-suite-based functional correctness, although non-smelly requirements could still produce faulty code. We also found that different smell categories had similar effects. These findings provide additional empirical evidence of the importance of requirement quality in LLM-assisted code generation, while showing that high-quality requirements alone do not guarantee correctness, as these depends on several factors, including the LLM. Compared with previous work, our results suggest that the impact of requirement smells depends on the software engineering task: whereas their effects on traceability were modest, code generation appears more sensitive. Overall, this work motivates further investigation into task-dependent quality effects in LLM-assisted software engineering.

发表机构

  • fortiss GmbH(fortiss有限公司)
  • Netlight Consulting(Netlight咨询公司)
  • Technical University of Munich(慕尼黑工业大学)
  • Chalmers University of Technology(查尔姆斯理工大学)
  • Trinity College Dublin(都柏林三一学院)
  • University of Duisburg-Essen(杜伊斯堡-埃森大学)
  • Blekinge Institute of Technology(布莱金厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑