提示工程技术在大语言模型版本间的老化现象
Aging of Prompt Engineering Techniques Across LLM Versions
浏览论文内容
中文总结 AI 辅助
本研究通过评估三类模型对的五种提示技术,发现提示工程效果随LLM代际以模型家族特定方式老化,需针对性调整而非照搬。
中文摘要 AI 辅助
提示工程及提示工程技术(PETs)已成为AI系统软件工程的重要组成部分。然而,新的大语言模型(LLMs)频繁发布,目前仍不清楚提示工程技术的有效性在大语言模型的连续代际间会发生怎样的变化。为此,我们对Khojah等人(2025年)的研究进行了部分重复。我们评估了五种技术——零样本(Zero-Shot)、少样本(Few-Shot)、思维链(Chain-of-Thought, CoT)、对比思维链(Contrastive Chain-of-Thought, CCoT),以及适配后的程序思维(Program-of-Thought, PoT)版本——在六个经指令微调的模型上的表现,这些模型分为三组版本对:GPT-3.5-Turbo/GPT-4o、Qwen2 7B Instruct/Qwen2.5 7B Instruct,以及Mistral-7B-Instruct/Mistral-Large。我们使用了CodePromptEval数据集的一个清理后的子集,包含218个上下文丰富的Python函数,共19620次生成结果,通过基于pass@k的功能正确性对模型对在函数级代码生成任务上进行评估。我们发现,提示工程会以特定于模型家族的方式“老化”:较新的GPT模型从结构化提示中获得的边际增益递减,甚至出现负增益,这表明模型内部越来越多地内化了遵循指令和推理的框架;而Qwen模型仍能从少样本和CCoT中获得显著收益;Mistral模型的表现则较为复杂,CCoT带来的收益持续存在,但CoT和PoT的收益有所减弱。我们的结果表明,有效的提示策略必须针对每个模型家族和代际进行调整,而非直接照搬。这为未来研究自适应的、感知模型的提示方法,以及更广泛的多维度代码质量评估提供了动力。
英文摘要
Prompt engineering and prompt engineering techniques (PETs) have become an integral part of software engineering for AI systems. However, new LLMs are released frequently and it remains unclear how the effectiveness of prompt engineering techniques changes across successive generations of Large Language Models (LLMs). To this end, we conduct a partial replication of the study by Khojah et al. (2025). We evaluate five techniques - Zero-Shot, Few-Shot, Chain-of-Thought (CoT), Contrastive Chain-of-Thought (CCoT), and an adapted version of Program-of-Thought (PoT) - on six instruction-tuned models grouped into three version pairs: GPT-3.5-Turbo/GPT-4o, Qwen2 7B Instruct/Qwen2.5 7B Instruct, and Mistral-7B-Instruct/Mistral-Large. We use a cleaned subset of the CodePromptEval dataset with 218 context-rich Python functions and 19,620 total generations assessed via pass@k-based functional correctness to evaluate model pairs on function-level code generation tasks. We show that prompt engineering "ages" in a model-family-specific way: Newer GPT models exhibit diminishing or even negative marginal gains from structured prompting, suggesting that instruction-following and reasoning scaffolds are increasingly internalized, whereas Qwen models continue to benefit substantially from Few-Shot and CCoT. Mistral models show mixed behavior with persistent gains from CCoT but attenuated benefits from CoT and PoT. Our results imply that effective prompting strategies must be adapted per model family and generation rather than transferred unchanged. This motivates future work on adaptive, model-aware prompting and broader, multi-dimensional code quality evaluation.