arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2411.02093cs.SE

先进语言模型是否消除了软件工程中对提示工程的需求?

Do Advanced Language Models Eliminate the Need for Prompt Engineering in Software Engineering?

Guoqing Wang, Zeyu Sun, Zhihao Gong, Sixiang Ye, Yizhou Chen, Yifan Zhao, Qingyuan Liang, Dan Hao

更新

AI总结:

本研究针对GPT-4o、o1等先进大语言模型,重新评估代码生成等三类软件工程任务中的提示工程效果,发现早期提示技术收益减弱甚至反作用,推理模型仅在复杂推理任务有优势,为模型与提示选择提供了实用指导。

AI中文摘要:

大语言模型(Large Language Models, LLMs)已显著推动了软件工程(software engineering, SE)任务的发展,提示工程技术则进一步提升了其在代码相关领域的性能。然而,基础大语言模型(如非推理模型GPT-4o和推理模型o1)的快速发展,引发了关于这些提示工程技术是否仍然有效的疑问。本文开展了一项广泛的实证研究,在这些先进大语言模型的语境下重新评估了多种提示工程技术。我们聚焦于三项代表性的软件工程任务,即代码生成、代码翻译和代码摘要,评估了以下问题:使用先进模型时提示工程技术是否仍能带来性能提升、推理模型相较于非推理模型的实际效果如何,以及使用这些先进模型的收益是否能抵消其增加的成本。研究结果表明,为早期大语言模型开发的提示工程技术应用于先进模型时,带来的收益可能会减弱,甚至会阻碍性能。在推理大语言模型中,其内置的复杂推理能力降低了复杂提示的作用,有时简单的零样本提示反而更有效。此外,尽管推理模型在需要复杂推理的任务上表现优于非推理模型,但在无需推理的任务上优势极小,还可能产生不必要的成本。基于本研究,我们为从业者提供了选择合适的提示工程技术和基础大语言模型的实用指导,需考虑任务需求、运营成本和环境影响等因素。我们的工作有助于更深入地理解如何在软件工程任务中有效利用先进大语言模型,为未来的研究和应用开发提供参考。

英文摘要:

Large Language Models (LLMs) have significantly advanced software engineering (SE) tasks, with prompt engineering techniques enhancing their performance in code-related areas. However, the rapid development of foundational LLMs such as the non-reasoning model GPT-4o and the reasoning model o1 raises questions about the continued effectiveness of these prompt engineering techniques. This paper presents an extensive empirical study that reevaluates various prompt engineering techniques within the context of these advanced LLMs. Focusing on three representative SE tasks, i.e., code generation, code translation, and code summarization, we assess whether prompt engineering techniques still yield improvements with advanced models, the actual effectiveness of reasoning models compared to non-reasoning models, and whether the benefits of using these advanced models justify their increased costs. Our findings reveal that prompt engineering techniques developed for earlier LLMs may provide diminished benefits or even hinder performance when applied to advanced models. In reasoning LLMs, the ability of sophisticated built-in reasoning reduces the impact of complex prompts, sometimes making simple zero-shot prompting more effective. Furthermore, while reasoning models outperform non-reasoning models in tasks requiring complex reasoning, they offer minimal advantages in tasks that do not need reasoning and may incur unnecessary costs. Based on our study, we provide practical guidance for practitioners on selecting appropriate prompt engineering techniques and foundational LLMs, considering factors such as task requirements, operational costs, and environmental impact. Our work contributes to a deeper understanding of effectively harnessing advanced LLMs in SE tasks, informing future research and application development.

↑