大语言模型代码生成中的复合提示约束:格式、角色与紧急性的析因研究
Compound Prompt Constraints in LLM Code Generation: A Factorial Study of Format, Persona, and Urgency
浏览论文内容
中文总结 AI 辅助
本研究通过3×3×3析因实验,探究格式、角色、紧急性三类复合提示约束对LLM代码生成可靠性的影响,发现架构依赖的超加性退化,提出复合提示测试应成为相关流程可靠性评估的标准环节。
中文摘要 AI 辅助
大型语言模型(LLM)正越来越多地用于软件工程流程中的代码生成,生产环境中的提示通常会结合多种约束条件。本文开展了一项全析因实证研究,探究输出格式、角色分配与紧急性框架如何共同影响LLM代码生成的可靠性。我们采用受控的3×3×3设计评估全部27种组合,并将每种复合条件分解为加性预测项与捕捉超加性退化的残差交互项。该研究使用HumanEval+的全部164个问题,覆盖GPT-4o系列、GPT-4.1系列及o3-mini共5款OpenAI模型,生成22140次贪婪解码评估。一个感知格式的提取管道将格式失败与推理失败分离开来,采用McNemar检验、优势比及95%置信区间评估显著性。结果显示,复合约束会产生架构依赖的退化,无法通过单因素实验预测;GPT-4o系列表现出一致的超加性效应,其pass@1降低幅度超出加性预测3-12个百分点,其中GPT-4o-mini在JSON+专家角色+中等紧急性组合上的交互效应最大,达-12.2个百分点;JSON组合的交互效应通常大于XML。相比之下,GPT-4.1系列基本不受影响,而o3-mini呈现出质的不同模式,结构化输出约束可提升性能。这些发现表明,脆弱性依赖于架构而非规模,单独中性或有益的约束组合可能导致严重退化,复合提示测试应成为LLM辅助工程流程可靠性评估的标准环节。
英文摘要
Large language models (LLMs) are increasingly used in software engineering pipelines for code generation, where production prompts often combine multiple constraints. This paper presents a full-factorial empirical study of how output formatting, persona assignment, and urgency framing jointly affect LLM code-generation reliability. We evaluate all 27 combinations in a controlled 3x3x3 design and decompose each compound condition into an additive prediction and a residual interaction term that captures super-additive degradation. The study uses all 164 HumanEval+ problems across five OpenAI models from the GPT-4o family, GPT-4.1 family, and o3-mini, yielding 22,140 greedy-decoding evaluations. A format-aware extraction pipeline separates formatting failures from reasoning failures, and significance is assessed with McNemar's test, odds ratios, and 95% confidence intervals. Results show that compound constraints can produce architecture-dependent degradation not predictable from single-factor experiments. The GPT-4o family exhibits consistent super-additive effects, with pass@1 reductions 3-12 percentage points beyond additive predictions; the largest interaction is -12.2 pp on GPT-4o-mini for JSON + expert persona + moderate urgency. JSON combinations generally produce larger interactions than XML. In contrast, the GPT-4.1 family is largely resistant, while o3-mini shows a qualitatively different pattern in which structured output constraints can improve performance. These findings show that vulnerability is architecture-dependent rather than size-dependent, that individually neutral or beneficial constraints can combine to cause substantial degradation, and that compound-prompt testing should be standard in reliability assessment for LLM-assisted engineering pipelines.
发表机构
- Embry-Riddle Aeronautical University(埃默里航空大学)
- Purdue University Northwest(普渡大学西北分校)
机构由 AI 辅助整理,请以论文原文为准。