AI 中文总结
本研究通过生成9000个语言受控的提示变体,利用5个开源LLM在需求分类任务上证实,提示的语言特征可显著预测LLM性能,为低成本提示选择提供了可解释的先验方法。
AI 中文摘要
背景:大语言模型(LLM)的输出对提示的表述高度敏感,措辞的微小变化会显著影响输出质量。这一问题在软件工程中尤为重要,因为提示指导需求分析、代码生成和工件合成,不当的提示会产生不可靠的工件,而从业者在推理前缺乏评估提示的原则性方法,导致提示选择依赖成本高昂的LLM调用和试错优化。目标:本研究探究提示的可测量语言属性是否能在推理前预测LLM性能,从而实现低成本的提示选择与优化,研究在针对F1、F2、精确率和召回率的二分类需求分类任务上进行验证。方法:从100个初始提示出发,通过改变30个语言指标生成9000个语言受控的提示变体,使用5个开源LLM在625个带标注的需求上进行评估;通过分层10折交叉验证结合基于置换的显著性检验训练回归预测器,特征重要性分析识别跨LLM和模型特定的预测因子。结果:语言特征对所有目标的提示性能均有显著预测作用,R²在[0.38,0.42]区间内,q值均小于0.05;句法和词法句法特征贡献了大部分预测信号,跨LLM的预测因子包括复合依存分布、连词密度以及词/句子长度,反映出模型对领域词汇和复杂结构的敏感性。结论:研究结果对提示工程具有实际意义,包括降低LLM性能的语言模式与增加人类理解难度的模式存在重叠,以及词汇多样性作为质量维度的无关性;更广泛而言,语言分析结合标准回归方法,能在成本高昂的优化流程前提供有效、可解释且低成本的先验判断。
英文摘要
Background. LLM outputs are highly sensitive to prompt formulation: small wording changes can substantially affect output quality. This matters in software engineering, where prompts guide requirements analysis, code generation, and artefact synthesis. Poor formulations yield unreliable artefacts, yet practitioners lack principled ways to assess a prompt before inference, making selection depend on costly LLM calls and trial-and-error refinement. Aims. We investigate whether measurable linguistic properties of prompts can predict LLM performance before inference, enabling low-cost prompt selection and refinement, validated on binary requirements classification targeting F1, F2, precision, and recall. Method. We generate 9,000 linguistically controlled prompt variants from 100 initial prompts by varying 30 linguistic metrics, evaluated with five open-source LLMs on 625 annotated requirements. Regression predictors are trained via stratified 10-fold cross-validation with permutation-based significance testing; feature importance analysis identifies cross-LLM and model-specific predictors. Results. Linguistic features significantly predict prompt performance across all targets (R2 in [0.38,0.42], q<0.05). Syntactic and morphosyntactic features drive most predictive signal; cross-LLM predictors include compound dependency distribution, conjunction density, and word/sentence length, reflecting sensitivity to domain vocabulary and complex structures. Conclusions. Results suggest practical implications for prompt engineering, including overlap between linguistic patterns that reduce LLM performance and those that increase human comprehension difficulty, and the irrelevance of lexical variety as a quality dimension. More broadly, linguistic profiling combined with standard regression provides an effective, interpretable, low-cost prior before costly optimisation pipelines.